The llama.cpp project has released version b11062, which includes a pull request enabling sparse Flash Attention (FA) support for the Qwen4 model.
This update is part of the broader b11062 release, which provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, GPU (CUDA, ROCm, Vulkan), and AI accelerator backends.
The specific change referenced in the release notes adds sparse FA capabilities for Qwen4 to improve efficiency.