The llama.cpp project released version b11203, which includes a change to the CUDA Fast Walsh-Hadamard Transform (FWHT) to accept F16 inputs. Previously, the CUDA FWHT only accepted F32 inputs; this update makes the source type a template parameter so the kernel reads an F16 source directly without requiring a converted copy.
The implementation ensures that `supports_op` and the dispatch call share a single predicate checking contiguity and shape conditions, closing a gap where mismatches could cause assertion failures. Test-backend-ops on an A10 GPU passed all 1297 MUL_MAT tests, including 6 new F16 Hadamard cases alongside existing F32 ones.
This release provides binaries for macOS, Linux, Windows, Android, and openEuler across CPU, CUDA, ROCm, Vulkan, OpenVINO, SYCL, and Snapdragon backends.