The llama.cpp b11216 release introduces Fast Walsh-Hadamard Transform (FWHT) kernels in the SYCL backend that support block widths greater than 512. Previously, larger widths fell back to a dense GEMM with O(n^2) complexity; this change enables O(n log n) performance for these sizes.
- The new `fwht_kernel_wide` implementation runs one row per work-group and uses local memory and registers to optimize butterfly operations.
- It supports block widths of 1024, 2048, 4096, and 8192 via the Kronecker/Paley construction.
- The release includes binaries for macOS, Linux, Windows, Android, and openEuler across CPU, GPU, and NPU backends.
- Verification was performed using a standalone harness on an Intel oneAPI DPC++ 2026.1 SYCL CPU device, confirming correctness against a reference implementation.
This update improves computational efficiency for large block sizes in SYCL environments while maintaining compatibility with existing hardware configurations.