The llama.cpp project has released version b11012, which introduces Vulkan support for qwen4exp hc operations via pull request #28988. This update also includes a fix for a stale comment.
The release provides binaries and frameworks across multiple platforms and hardware accelerators:
- macOS: Apple Silicon (arm64), Intel (x64), and iOS XCFramework.
- Linux: Ubuntu builds for CPU, Vulkan, CUDA 12/13, ROCm 10.0, OpenVINO, and SYCL.
- Windows: Builds for CPU, OpenCL Adreno, CUDA 12/13, Vulkan, OpenVINO, SYCL, and ROCm 10.0.
- Android: arm64 CPU support.
This release enables users to run qwen4exp models with high-compute operations on Vulkan-compatible GPUs across supported operating systems.