The llama.cpp project has released version b10505, which introduces a new `dedup-cache-models` preset option for the server component.

  • The release includes binaries for macOS (Apple Silicon and Intel), Linux (Ubuntu x64, arm64, s390x, Vulkan, OpenVINO, SYCL), Android, Windows (CPU, CUDA 12/13, Vulkan, OpenVINO, SYCL, ROCm), and openEuler.
  • macOS Apple Silicon KleidiAI support is disabled in this build.
  • Ubuntu x64 ROCm 7.14 support is also disabled.
  • A new UI package is available for download alongside the platform-specific binaries.