The llama.cpp project released build b10759, which includes a fix to avoid KleidiAI buffer type initialization on dispatch.

  • The release provides binaries for macOS (Apple Silicon and Intel), iOS, Linux (Ubuntu x64, arm64, s390x), Windows (CPU, CUDA 12/13, Vulkan, OpenCL, ROCm, SYCL, OpenVINO), Android, and openEuler.
  • KleidiAI support is explicitly disabled for macOS Apple Silicon in this build.
  • Previews for Windows arm64 with CUDA 13 are included alongside standard CUDA 12.4 and 13.3 releases.