The llama.cpp b11413 release introduces a new Vulkan implementation of sparse flash attention that supports quantized Key and Value tensors. This update addresses performance issues in the previous compaction method, which was too costly for decoding tasks with large context windows.
- The new approach splits rows into contiguous segments processed by subgroups or threads, using ballot counting over coalesced loads to maintain an ascending index list.
- Pre-built binaries are available for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA 12/13, ROCm, OpenVINO, SYCL, Snapdragon), Windows (CPU, Vulkan, CUDA, OpenVINO, SYCL, ROCm), Android, and iOS.
- The release also includes the llama.cpp UI and attestations for security verification.