The llama.cpp project has released version b11097, which introduces an OpenCL binary kernel for the A8 Q4_0 non-MoE dp4a format.

  • The update adds support for the specific A8 Q4_0 quantization format via a new dp4a kernel.
  • Binaries are provided for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA 12/13, ROCm 10.0, OpenVINO, SYCL), Windows (CPU, Vulkan, CUDA 12/13, ROCm 10.0, OpenVINO, SYCL), Android, and openEuler.
  • An iOS XCFramework and a standalone UI build are also included in the release assets.