The llama.cpp project released version b11206, which introduces support for the Hexagon backend sampler. This update includes several changes to the sampling pipeline and operations on the Hexagon NPU.

  • Added STEP and SUM operations to the hex-sampler.
  • Updated CPY to support sampling cases.
  • Added chunking support in hex-binary to handle large logits.
  • Implemented a basic version of ARGMAX for hex-argmax.
  • Fixed indexing issues for dim 1 broadcasts across dim 2 slices.
  • Disabled the autovectorizer to use explicit hints for critical loops.

The release provides binaries for macOS, Linux, Windows, Android, and openEuler across various hardware backends including CPU, GPU, and NPU.