The llama.cpp project released version b11440, which includes a fix for speculative MTP initialization. The update ensures the scheduler is re-reserved when NextN extraction flags change, preventing issues with unmasked extraction.

  • Speculative MTP init now enables NextN extraction on target and draft contexts after creation.
  • The trunk graph keeps every token through the last layer instead of cropping to output rows.
  • Invalidating the reserve when flags change allows the next compute to re-reserve with the new graph shape.
  • Binaries are available for macOS, Linux, Windows, Android, and iOS across CPU, GPU, and NPU backends.

This change prevents GGML_SCHED_DEBUG_REALLOC errors by ensuring correct batch shape allocation during decoding.