The llama.cpp b10236 release introduces a Metal implementation of the DSv4 Lightning Indexer, specifically supporting 128-dimensional inputs with F32 queries and weights alongside F16 keys and masks.
- The update includes tiled and tail kernels for handling KV lengths near 8- and 64-element boundaries.
- K tiles are staged and dequantized in F16 threadgroup memory before simdgroup matrix loads, supporting F32, F16, BF16, and various quantization formats.
- Benchmarks show improved prefill throughput (pp512) at sequence lengths of 10k, 20k, and 30k tokens compared to previous versions.
This optimization enhances inference performance on Apple Silicon devices by improving memory access patterns for the Lightning Indexer.