The llama.cpp project has released version b10213, introducing support for rotated key-value cache quantization. This update is available across a wide range of platforms and hardware accelerators.
- macOS Apple Silicon (arm64) and Intel (x64)
- iOS via XCFramework
- Linux Ubuntu x64 and arm64 (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL FP32/FP16)
- Windows x64 and arm64 (CPU, OpenCL Adreno, CUDA 12/13, Vulkan, OpenVINO, SYCL, HIP)
- Android arm64 (CPU)
- openEuler x86 and aarch64 (ACL Graph)
The release includes binaries for the main application as well as a separate UI package.