The llama.cpp project has released version b10213, introducing support for rotated key-value cache quantization. This update is available across a wide range of platforms and hardware accelerators.

  • macOS Apple Silicon (arm64) and Intel (x64)
  • iOS via XCFramework
  • Linux Ubuntu x64 and arm64 (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL FP32/FP16)
  • Windows x64 and arm64 (CPU, OpenCL Adreno, CUDA 12/13, Vulkan, OpenVINO, SYCL, HIP)
  • Android arm64 (CPU)
  • openEuler x86 and aarch64 (ACL Graph)

The release includes binaries for the main application as well as a separate UI package.