The llama.cpp project released build b11272, which includes optimizations for the Hexagon backend. The primary change focuses on optimizing the concatenation operation to improve performance on Qualcomm Snapdragon hardware.

  • Reduced packets in the gather/transpose hot loop by gathering directly into the destination buffer and using a special instruction for gather sync.
  • Replaced software division calls with fastdiv for faster computation.
  • Optimized the DMA-HVX pipeline and added transpose helpers for the Hexagon concat operation.

This update provides binaries for macOS, Linux, Windows, Android, and openEuler across various CPU, GPU, and NPU backends.