The llama.cpp project released build b11207, which introduces support for tiled Q4_0 and Q8_0 GET_ROWS operations on the Hexagon backend. This update also includes fixes to macros, register spills, and the DMA pipeline, alongside improvements to kernel selection logic and re-enabled vectorization.
- Added tiled Q4_0 and Q8_0 GET_ROWS support for Hexagon using tiled HVX dequantization.
- Fixed macros and register spills in hex-get-rows while cleaning up checks for unsupported operations.
- Improved the DMA pipeline and simplified kernel selection logic.
- Re-enabled the vectorizer after a regression was identified during a sampler update.
The release provides binaries for macOS, Linux, Windows, Android, and openEuler across various hardware backends including CPU, GPU, NPU, CUDA, ROCm, Vulkan, OpenVINO, and SYCL.