The llama.cpp project has released version b11280, which includes a change to reduce tensor allreduce synchronization by utilizing pinned host buffers.
This update is available across multiple platforms and backends, including macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA 12/13, ROCm 10.0, OpenVINO, SYCL, and Snapdragon), Windows (CPU, OpenCL, CUDA 12/13, Vulkan, OpenVINO, SYCL, and ROCm 10.0), Android, and iOS.
The release also provides binaries for the llama.cpp UI.