The llama.cpp project released version b11335, which includes a CUDA optimization for Volta GPUs. The update routes the sm_70 architecture to the existing Turing MMVQ nwarps table instead of falling through to the generic configuration.

  • Volta (sm_70) now uses the TURING tuning for K-quant batch-1 decode, changing the warp count from 4 to 2.
  • Benchmarks on a Tesla V100 with Qwen3.8-27B show a +3.17% throughput increase (+1.091 t/s) with bit-identical perplexity.
  • The change is derived from the anyei/llamacpp-v100 fork and applies to both device and host table selectors.

This adjustment improves inference speed on Volta hardware without affecting model accuracy.