The llama.cpp project released build b11338, introducing a shared strided Direct Memory Access (DMA) copy mechanism for the CPY and CONCAT operations. This update enables any-dimension CONCAT via DMA on Hexagon architectures.
- Implemented shared strided DMA copy for CPY and CONCAT operations.
- Enabled any-dimension CONCAT via DMA.
- Removed broken CONCAT_DMA_MIN_ROW logic that was incompatible with 64-bit DMA.
- Added missing dma_queue_flush() calls and additional guards for unsupported conditions.
This release provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, GPU, and NPU backends.