The llama.cpp project released build b11362, which introduces a new tensor API flash attention kernel supporting F16 key-value (KV) storage. This update is available across macOS, Linux, Windows, Android, and iOS platforms with various backend options including CUDA, Vulkan, and ROCm.
- Adds tensor FA kernels for specific dimensions: DK=DV=512, DK=576/DV=512, and DK=192/DV=128.
- The new kernel supports attention sinks, ALiBi, and logit softcap mechanisms.
- Pre-built binaries are provided for CPU, GPU (CUDA 12/13, Vulkan, ROCm 10.0), and NPU (Snapdragon Hexagon) backends.
This release enables more efficient attention computation on supported hardware configurations.