The llama.cpp project released version b11160, introducing an int8 cooperative matrix (coopmat1) implementation for Vulkan on AMD RDNA3 and RDNA4 architectures. This update includes a dedicated shader that probes coopmat values directly to improve performance.

  • Adds IQ4_XS quantization support to the mul_mmq_cm1 integer matmul shader.
  • Implements inline scale application, scalar sums, and double buffering for the Vulkan backend.
  • Supports q8_0, q4_1, q5_0, q5_1, iq4_nl, mxfp4, q3_k, q4_k, q5_k, q6_k, and nvfp4 quantizations.
  • Includes binaries for macOS, Linux, Windows, Android, and openEuler across CPU, CUDA, ROCm, OpenVINO, and SYCL backends.