The llama.cpp project has released build b10603, which introduces Multi-Token Prediction (MTP) support for the GLM-4.5-Air model.

  • Adds MTP capability for GLM-4.5-Air via pull request #26534.
  • Provides binaries for macOS (Apple Silicon and Intel), iOS, Linux (CPU, Vulkan, ROCm, OpenVINO, SYCL), Android, Windows (CPU, CUDA 12/13, Vulkan, OpenCL, OpenVINO, SYCL, ROCm), and openEuler.

This release enables users to run GLM-4.5-Air with MTP on a wide range of hardware architectures.