The llama.cpp project released version b11440, which includes a fix for speculative MTP initialization. The update ensures the scheduler is re-reserved when NextN extraction flags change, preventing issues with unmasked extraction.
- Speculative MTP init now enables NextN extraction on target and draft contexts after creation.
- The trunk graph keeps every token through the last layer instead of cropping to output rows.
- Invalidating the reserve when flags change allows the next compute to re-reserve with the new graph shape.
- Binaries are available for macOS, Linux, Windows, Android, and iOS across CPU, GPU, and NPU backends.
This change prevents GGML_SCHED_DEBUG_REALLOC errors by ensuring correct batch shape allocation during decoding.