The llama.cpp project released version b11206, which introduces support for the Hexagon backend sampler. This update includes several changes to the sampling pipeline and operations on the Hexagon NPU.
- Added STEP and SUM operations to the hex-sampler.
- Updated CPY to support sampling cases.
- Added chunking support in hex-binary to handle large logits.
- Implemented a basic version of ARGMAX for hex-argmax.
- Fixed indexing issues for dim 1 broadcasts across dim 2 slices.
- Disabled the autovectorizer to use explicit hints for critical loops.
The release provides binaries for macOS, Linux, Windows, Android, and openEuler across various hardware backends including CPU, GPU, and NPU.