The llama.cpp project has released version b10425, which includes a port of the SYCL optimization to fuse the gated-delta-net state writeback copy.
- The change improves performance by approximately 1.2% on Arc Pro B70 hardware when running Qwen 3.6 27B with interleaved A/B passes.
- Pre-fill latency (pp2048) remains flat, while token generation speed increases slightly.
- Binaries are available for macOS, Linux, Windows, Android, and openEuler across CPU, CUDA, Vulkan, ROCm, OpenVINO, and SYCL backends.
This release provides updated builds for various platforms and accelerators to support the latest model optimizations.