The llama.cpp project released build b11232, which includes updates to the ggml-cpu backend. The release focuses on enabling tiled flash attention for non-vector-multiple head dimensions on x86 architectures.
- Enabled tiled flash attention for non-vector-multiple head dims on x86.
- Added AVX2 support for masked loading and storing in simd_gemm_ukernel_tail.
- Fixed FA softcap handling for padded KV tiles.
This update provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, GPU, and NPU backends.