The llama.cpp project released build b11338, introducing a shared strided Direct Memory Access (DMA) copy mechanism for the CPY and CONCAT operations. This update enables any-dimension CONCAT via DMA on Hexagon architectures.

  • Implemented shared strided DMA copy for CPY and CONCAT operations.
  • Enabled any-dimension CONCAT via DMA.
  • Removed broken CONCAT_DMA_MIN_ROW logic that was incompatible with 64-bit DMA.
  • Added missing dma_queue_flush() calls and additional guards for unsupported conditions.

This release provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, GPU, and NPU backends.