The llama.cpp project has released version b11214, which includes a key update to the HIP backend. This release enables the fattn-mma kernel on CDNA architectures when dkq is greater than 256, specifically targeting performance improvements for large batch sizes.
- The HIP backend now supports the fattn-mma kernel for dkq > 256 on CDNA hardware.
- CI checks for hip-quality-check have been updated to ignore spills for very large mfma mma kernels.
- Binaries are provided for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA, ROCm, OpenVINO, SYCL, Snapdragon), Windows (CPU, CUDA, Vulkan, OpenVINO, SYCL, ROCm), Android, and openEuler.
This update allows users running llama.cpp on AMD CDNA GPUs to utilize optimized attention kernels for larger batch sizes.