The llama.cpp project has released build b11382, which includes a key update to its WebGPU backend. This release adds 16-bit floating-point (f16) support specifically for the `fill` and `set_rows` operations within the WebGPU implementation.
- The source code change is tracked in pull request #29897 on the ggml-org/llama.cpp repository.
- Binaries are provided for macOS (Apple Silicon and Intel), iOS, Linux (CPU, Vulkan, CUDA 12/13, ROCm, OpenVINO, SYCL, Snapdragon), Windows (CPU, OpenCL, CUDA, Vulkan, OpenVINO, SYCL, ROCm), Android, and openEuler.
- The release also includes the standard llama.cpp UI binary.