The llama.cpp project released build b11362, which introduces a new tensor API flash attention kernel supporting F16 key-value (KV) storage. This update is available across macOS, Linux, Windows, Android, and iOS platforms with various backend options including CUDA, Vulkan, and ROCm.

  • Adds tensor FA kernels for specific dimensions: DK=DV=512, DK=576/DV=512, and DK=192/DV=128.
  • The new kernel supports attention sinks, ALiBi, and logit softcap mechanisms.
  • Pre-built binaries are provided for CPU, GPU (CUDA 12/13, Vulkan, ROCm 10.0), and NPU (Snapdragon Hexagon) backends.

This release enables more efficient attention computation on supported hardware configurations.