The llama.cpp project released version b11195, introducing a tiled matrix multiplication implementation for k-quants in the CPU backend. This change unpacks quantized data into 256x256 int8 tiles and computes results using microkernels to improve performance on large matrix operations.
- Tiled mul_mat provides a 3-6x speed improvement for large matrix multiplications, with break-even at 4096x64 * 64x4096 dimensions.
- Error rates remain trivial, with maximum errors around 1-e04 and RMSE around 1-e05.
- The release includes fixes for ARM and Windows builds, unified IQP support for Q5_K and IQ4_XS benchmarks, and optimized AVX2 kernels.
The update improves inference speed for large matrix operations while maintaining numerical accuracy across supported platforms.