The llama.cpp project has released version b10336, which includes a refactor of several wgsl files and simplification of flash_attn wgsl within the ggml-webgpu component.
This update provides binaries for macOS (Apple Silicon and Intel), iOS, Linux (Ubuntu with CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL), Android, Windows (CPU, OpenCL, CUDA 12/13, Vulkan, OpenVINO, SYCL, HIP), and openEuler platforms.
The release also includes the standard UI build for users to interact with the model inference engine.