The llama.cpp project has released build b11333, which introduces bfloat16 support for the MUL_MAT, MUL_MAT_ID, and GET_ROWS operations within its WebGPU backend.

This update is included in precompiled binaries for macOS (Apple Silicon and Intel), iOS, Linux (Ubuntu x64, arm64, s390x), Windows, Android, and openEuler. The release provides builds for various hardware accelerators including CUDA 12 and 13, Vulkan, ROCm 10.0, OpenVINO, SYCL, and Snapdragon NPUs.

The addition of bfloat16 precision to these specific WebGPU matrix operations allows for more efficient computation on compatible GPUs.