The llama.cpp project has released version b11424, which includes a fix for a shared memory write-out-of-bounds error in the Vulkan backend when using Flash Attention.

This update provides precompiled binaries for macOS (Apple Silicon and Intel), Linux (CPU, CUDA, ROCm, OpenVINO, SYCL, Snapdragon), Windows (CPU, CUDA, Vulkan, OpenVINO, SYCL, ROCm), Android, and iOS. The release also includes the llama.cpp UI.

The build for KleidiAI on macOS Apple Silicon is currently disabled in this version.