The llama.cpp project released build b11307, which fixes layer input ordering to preserve the original batch sequence during speculative decoding. This change ensures that token order is correctly maintained for unmasked NextN embeddings and compatible with tensor splits.
- Restores original row order after synchronization by copying microbatch tensors from offset zero.
- Uses original-token mapping for unmasked NextN rows, even when layer-input capture is disabled.
- Fixes WebGPU reservation and OpenVINO hidden-state capture issues.
- Validates across 256 CPU/CUDA/tensor configurations with Qwen3.8-27B completing MT-Bench at concurrency 16.
The update ensures correctness in speculative decoding workflows by preventing order corruption in layer inputs and embeddings.