The llama.cpp project released build b11254, which includes a fix for how input tensors are collected in the computation graph. Previously, `graph_inputs` was populated during graph splitting, causing issues with pipeline parallelism where switching between batches consuming different inputs led to spurious backend ID comparisons and forced scheduler re-reservations.
- The change collects inputs after the split from all input leafs of the graph, ensuring composition depends only on which inputs exist rather than which are used as sources.
- This prevents the scheduler from recording smaller input sizes (e.g., `n_outputs = 0`) and aborting later via `GGML_SCHED_DEBUG_REALLOC`.
- The release provides binaries for macOS (Apple Silicon and Intel), iOS, Linux (CPU, Vulkan, CUDA 12/13, ROCm, OpenVINO, SYCL, Snapdragon), Android, Windows (CPU, OpenCL, CUDA, Vulkan, OpenVINO, SYCL, ROCm), and openEuler.
This fix stabilizes execution in pipeline parallelism scenarios by eliminating unnecessary graph recompositions caused by input order changes.