The llama.cpp project has expanded its generic few-row Matrix Multiply (MMA) kernel to support additional source data types, including BF16, Q1_0, Q2_0, MXFP4, Q2_K, Q3_K, TQ2_0, and various IQ types. This update allows the optimized kernel to handle any type with a 16-weight dequantizer, applying specific row count thresholds on M3 Ultra hardware where performance begins to exceed current kernels.
- The new support covers BF16, Q1_0, Q2_0, MXFP4, Q2_K, Q3_K, TQ2_0, and IQ types.
- Performance thresholds vary by type: 5 rows for TQ2_0, 4 for BF16, 3 for MXFP4/Q2_0/Q2_K/IQ4_NL, and 2 for others.
- Benchmark results on M3 Ultra show performance ratios ranging from 0.23 to 1.01 depending on row count and kernel overlap.
This change improves inference efficiency for a wider range of quantization formats by leveraging the MMA kernel's capabilities at low row counts.