The llama.cpp project has released version b11460, which introduces support for the IQ3_S multi-column MMVQ quantization format within its SYCL backend. This update is included in the new build alongside standard CPU and GPU acceleration options.

  • The release adds IQ3_S multi-column MMVQ support to the SYCL implementation via pull request #29500.
  • Binaries are provided for macOS (Apple Silicon and Intel), Linux (x64, arm64, s390x), Windows, Android, and openEuler.
  • GPU acceleration is available via CUDA 12/13, Vulkan, ROCm 10.0, OpenVINO, and OpenCL Adreno across supported platforms.
  • Specific builds are included for Snapdragon devices on Linux and Android, supporting CPU, Adreno GPU, and Hexagon NPU.

This release allows users to access the latest SYCL optimizations and broadens hardware compatibility for running llama.cpp models on diverse architectures.