The llama.cpp project released build b11227, which includes performance optimizations specifically targeting the handling of the `causal_attn` flag during multimodal processing.
- The scheduler no longer re-reserves memory when toggling `causal_attn`, eliminating expensive passes that previously slowed down vision inputs for models like Gemma.
- The qwen4exp indexer bias shape is now independent of `causal_attn`, preventing buffer reallocation failures under `GGML_SCHED_NO_REALLOC` constraints.
- Benchmarks on llama-server with gemma-4-26B-A4B show significant latency reductions, such as a 7.61x speedup for 24 images on RTX 4090 and up to 3.40x on H200 with large context settings.
These changes improve inference speed for multi-image and video inputs without altering the generated output or graph rebuilding behavior.