The llama.cpp project released build b11207, which introduces support for tiled Q4_0 and Q8_0 GET_ROWS operations on the Hexagon backend. This update also includes fixes to macros, register spills, and the DMA pipeline, alongside improvements to kernel selection logic and re-enabled vectorization.

  • Added tiled Q4_0 and Q8_0 GET_ROWS support for Hexagon using tiled HVX dequantization.
  • Fixed macros and register spills in hex-get-rows while cleaning up checks for unsupported operations.
  • Improved the DMA pipeline and simplified kernel selection logic.
  • Re-enabled the vectorizer after a regression was identified during a sampler update.

The release provides binaries for macOS, Linux, Windows, Android, and openEuler across various hardware backends including CPU, GPU, NPU, CUDA, ROCm, Vulkan, OpenVINO, and SYCL.