The llama.cpp project released version b11309, which includes a key optimization for the Hexagon backend. This update adds support for safe scatter mode to the ALLREDUCE operation, aiming to improve performance and stability on Qualcomm Snapdragon devices.

  • Optimizes Hexagon ALLREDUCE by adding safe scatter mode support.
  • Rewrites the hex-allreduce implementation to remove register spills.
  • Reduces excessive comments in the hex-allreduce code.
  • Provides binaries for macOS (Apple Silicon and Intel), Linux, Windows, Android, and openEuler across CPU, GPU, and NPU backends.

This release enables more efficient distributed inference on mobile hardware with Hexagon NPUs while maintaining broad compatibility across major desktop and mobile operating systems.