The llama.cpp project released build b11306, which addresses a crash caused by graph shape changes during decoding. The update modifies the handling of the `causal_attn` flag to prevent reallocation failures under specific scheduling constraints.
- After device decode, `causal_attn` is flipped off for half the ubatch and then back on for the remainder to ensure consistent node counts.
- This change prevents the graph from reallocating at an unchanged size, which previously aborted execution when `GGML_SCHED_NO_REALLOC` was active.
- The fix applies to decode architectures but is skipped for encode architectures.
The release includes binaries for macOS (Apple Silicon and Intel), iOS, Linux (CPU, CUDA 12/13, ROCm 10.0, OpenVINO, SYCL, Snapdragon), Android, Windows (CPU, CUDA 12/13, Vulkan, OpenVINO, SYCL, ROCm 10.0), and openEuler.