The llama.cpp project has added initial support for the GLM-5.3-Flash (GLM5-Next) model architecture, including migration to hybrid-index memory and early Multi-Token Prediction (MTP) capabilities.
- Initial MTP support was added, though stripped for the initial pull request to allow focused review.
- The k-pool layout is now kept across ubatches to prevent staleness after sequence edits and shared teardowns.
- Recurrent rollback checkpoints are correctly written for conv state and delta net state.
- A bug in build_attn_mha stream stride was fixed for non-contiguous queries, ensuring correct multi-stream prefill for GLM5-Next.
- Quantization protection filters were cleaned up to remove duplicate entries.
These changes ensure stable inference and correct state management for GLM5-Next models across various hardware backends including CUDA, ROCm, and CPU.