The llama.cpp project released version b11433, which introduces support for pooling operations on the Hexagon NPU. This update includes implementations for both 2D and 1D pooling layers.

  • Added pool_2d and pool_1d support for Hexagon backends.
  • Optimized HTP pooling boundaries and DMA pipelining.
  • Rewrote the DMA pipeline and added pool chunking support.
  • Removed vectorized scalar paths and simplified the chunk solver.

This release provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, GPU, and NPU architectures.