The llama.cpp project has released version b10827, which includes a fix to properly choose the weights pack for q4_K and q5_K matrix multiplication operations using OpenCL.

  • The release provides binaries for macOS (Apple Silicon and Intel), Linux (Ubuntu x64, arm64, s390x), Windows (CPU, CUDA 12/13, Vulkan, ROCm, SYCL), Android, and iOS.
  • OpenCL Adreno support is available on Windows arm64, while KleidiAI remains disabled for macOS Apple Silicon.
  • The update also includes a new UI build and attestation links for verification.