The llama.cpp project has released version b11012, which introduces Vulkan support for qwen4exp hc operations via pull request #28988. This update also includes a fix for a stale comment.

The release provides binaries and frameworks across multiple platforms and hardware accelerators:

  • macOS: Apple Silicon (arm64), Intel (x64), and iOS XCFramework.
  • Linux: Ubuntu builds for CPU, Vulkan, CUDA 12/13, ROCm 10.0, OpenVINO, and SYCL.
  • Windows: Builds for CPU, OpenCL Adreno, CUDA 12/13, Vulkan, OpenVINO, SYCL, and ROCm 10.0.
  • Android: arm64 CPU support.

This release enables users to run qwen4exp models with high-compute operations on Vulkan-compatible GPUs across supported operating systems.