The llama.cpp project released build b11017, which includes a Vulkan optimization to skip unnecessary Mixture of Experts (MoE) work in the mul_mm coopmat1 path.
This update provides pre-built binaries for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA 12/13, ROCm, OpenVINO, SYCL), Windows (CPU, OpenCL Adreno, CUDA 12/13, Vulkan, OpenVINO, SYCL, ROCm), Android, and openEuler.
The release also includes the llama.cpp UI.