The llama.cpp project has released build b10889, which includes a memory optimization to avoid allocating the V cache for the indexer since it is not used.

This update provides pre-built binaries for macOS (Apple Silicon and Intel), iOS, Linux (Ubuntu x64, arm64, s390x with CPU, Vulkan, ROCm 10.0, OpenVINO, and SYCL backends), Android (arm64), Windows (CPU, OpenCL Adreno, CUDA 12/13, Vulkan, OpenVINO, SYCL, ROCm 10.0), and openEuler.

The release also includes the llama.cpp UI.