The llama.cpp project released build b11375, which addresses a critical issue where non-contiguous ubatch cells forced graph reallocations that would abort when the GGML_SCHED_NO_REALLOC flag was set.

  • The fix gathers recurrent states once so the reserve covers every split, ensuring the worst-case reserve size is sufficient.
  • A single get_rows now gathers n_rs states, with ubatch and extra states treated as views of this single allocation.
  • Custom getters like mamba ssm_scan gather from the second state to avoid copying state for single sequence ubatches.
  • Views are built once per graph to keep host overhead unchanged.

This change ensures stable execution in constrained scheduling environments by preventing aborts caused by memory reallocation constraints.