The llama.cpp b11463 release includes a SYCL update that enables MKL flash attention for GLM-4.7 Flash's specific Multi-Latent Attention (MLA) shape, which was previously rejected by the dispatcher.

  • The change admits the 576/576/512 GQA-20 F16 shape to the MKL pipeline and handles the V cache as a strided view of K rows.
  • An attempt to store MKL flash attention scores in F16 to reduce memory traffic was included but subsequently reverted in this build.
  • On an Intel Arc Pro B70, the MLA optimization improved prompt processing (pp8192) from 432.80 to 1292.29 tok/s, a 2.99x speedup.
  • The release provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, CUDA, ROCm, Vulkan, OpenVINO, and SYCL backends.