The llama.cpp b11405 release updates the CUDA backend by tiling the lightning indexer over keys and tokens for 4 heads, addressing performance issues with small head counts.

  • The kernel stages queries and weights for all heads at once, widening each half-precision key element once to feed all heads without a float round trip.
  • Queries remain in float within shared memory to prevent overflow from half2 products when q * k exceeds the f16 range.
  • Each thread now scores two keys for a single token, reducing shared reads and keeping gfx908 at 63 VGPRs with no register spill.
  • This optimization makes the kernel 36x faster on an R9700 and slightly faster on CUDA hardware.

The release provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, GPU (CUDA, ROCm, Vulkan, OpenCL), and NPU backends.