The llama.cpp project has added initial support for the GLM-5.3-Flash (GLM5-Next) model architecture, including migration to hybrid-index memory and early Multi-Token Prediction (MTP) capabilities.

  • Initial MTP support was added, though stripped for the initial pull request to allow focused review.
  • The k-pool layout is now kept across ubatches to prevent staleness after sequence edits and shared teardowns.
  • Recurrent rollback checkpoints are correctly written for conv state and delta net state.
  • A bug in build_attn_mha stream stride was fixed for non-contiguous queries, ensuring correct multi-stream prefill for GLM5-Next.
  • Quantization protection filters were cleaned up to remove duplicate entries.

These changes ensure stable inference and correct state management for GLM5-Next models across various hardware backends including CUDA, ROCm, and CPU.