The llama.cpp project released build b11232, which includes updates to the ggml-cpu backend. The release focuses on enabling tiled flash attention for non-vector-multiple head dimensions on x86 architectures.

  • Enabled tiled flash attention for non-vector-multiple head dims on x86.
  • Added AVX2 support for masked loading and storing in simd_gemm_ukernel_tail.
  • Fixed FA softcap handling for padded KV tiles.

This update provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, GPU, and NPU backends.