The llama.cpp project released build b11375, which addresses a critical issue where non-contiguous ubatch cells forced graph reallocations that would abort when the GGML_SCHED_NO_REALLOC flag was set.
- The fix gathers recurrent states once so the reserve covers every split, ensuring the worst-case reserve size is sufficient.
- A single get_rows now gathers n_rs states, with ubatch and extra states treated as views of this single allocation.
- Custom getters like mamba ssm_scan gather from the second state to avoid copying state for single sequence ubatches.
- Views are built once per graph to keep host overhead unchanged.
This change ensures stable execution in constrained scheduling environments by preventing aborts caused by memory reallocation constraints.