The llama.cpp project released version b11403, which includes a CUDA optimization to use the MMVF kernel for thin f16 and bf16 matrix multiplications at small batch sizes.

  • Adjusted kernel selection logic for improved performance on NVIDIA GPUs.
  • Provided binaries for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA 12/13, ROCm, OpenVINO, SYCL, Snapdragon), Windows (CPU, OpenCL, CUDA 12/13, Vulkan, OpenVINO, SYCL, ROCm), Android, and openEuler.
  • Included iOS XCFramework and UI builds alongside the standard binaries.