The llama.cpp project has released version b10202, introducing a SYCL optimization that fuses the RMS_NORM and MUL operations.
This update provides pre-built binaries for macOS (Apple Silicon and Intel), iOS, Linux (Ubuntu x64, arm64, s390x), Windows, and Android. It supports various hardware backends including CUDA 12/13, Vulkan, ROCm 7.2, OpenVINO, SYCL FP32/FP16, HIP, and OpenCL Adreno.
The release also includes the standard llama.cpp UI and notes that KleidiAI support for macOS Apple Silicon is currently disabled.