The llama.cpp project has released build b10251, which introduces Multi-Token Prediction (MTP) support for the GLM-4.7-Flash model.

This release includes binaries for macOS (Apple Silicon and Intel), iOS, Linux (Ubuntu x64, arm64, s390x with CPU, Vulkan, ROCm 7.2, OpenVINO, and SYCL backends), Android (arm64), Windows (CPU, OpenCL Adreno, CUDA 12/13, Vulkan, OpenVINO, SYCL, HIP), and openEuler (ACL Graph).

The macOS Apple Silicon KleidiAI build has been disabled in this version.