The llama.cpp project released build b11184, introducing new Fast Walsh-Hadamard Transform (FWHT) kernels for the Metal backend that support block widths greater than 512. Previously, the Metal FWHT implementation was limited to widths between 64 and 512.
- The new kernel handles widths from 1024 through 8192 by running one row per threadgroup with 256 threads, allowing wider blocks to utilize threadgroup memory instead of relying solely on registers.
- Widths 64 to 512 continue to use the existing simdgroup kernel, while the new wide kernels support both F32 and F16 source types.
- The implementation includes size checks against device limits to prevent aborts on devices with insufficient threadgroup memory, allocating float[N] of memory which reaches 32 KB at width 8192.
- Backend operation tests on an M5 Pro device passed all checks for MUL_MAT_HADAMARD and MUL_MAT.
This update enables llama.cpp to perform Hadamard product operations on larger block sizes within the Metal backend, extending the range of supported tensor dimensions for Apple Silicon hardware.