llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API for mixed token and embedding inputs, along with support for the GLM-5.3-Flash 320B hybrid model and the Clef decision model.
- The new llama_batch_ext API supports per-token state embeddings for MTP and deepstack models.
- Qwen4Exp now offers high-quality support with MTP speculative decoding, providing approximately 1.5x decode speedup on DGX Spark.
- A new /v1/systemone server API supports five decision models: laya, julia-1, lev, openjev, and kev.
- Metal gains a tensor-API flash attention kernel for F16 KV and MMA mat-mul kernels up to 3x faster on Apple GPUs.
- The Web UI receives an overhaul with a Hugging Face Hub data layer and model download pipeline.
This release expands llama.cpp's capabilities for multimodal processing, speculative decoding performance, and decision model integration.