The Metal backend introduces new matrix multiplication kernels for 2 to 16 source rows, enabling faster speculative decoding on Apple Silicon devices lacking the tensor API. These kernels dequantize weights once and split K dimensions across threadgroups, offering significant performance gains over previous mat-vec approaches.

  • Implements specific MMA kernels for Q4_0, Q8_0, Q5_K, and generic paths for F32, F16, and other quantization types.
  • Activates these kernels on MTLGPUFamilyApple7+ hardware when row counts exceed thresholds where they outperform mat-vec kernels.
  • Adds fusion logic to combine MUL_MAT + ADD operations and batch up to 16 adjacent same-layout f32 copies into a single dispatch.
  • Updates the graph optimizer and encoder to handle device properties, concurrency checks, and view tracking for fused groups.

This optimization resolves performance regressions in speculative decoding on M3 Ultra hardware by ensuring that tensor operations remain efficient even with small batch sizes.