The llama.cpp b11372 release introduces optimizations for the qwen4exp model's lightning indexer, primarily halving the memory required for indexer scores during long context processing. This is achieved by computing head scores in place rather than materializing separate tensors for each head.

  • The qwen4exp indexer now reuses allocator buffers and computes scores via ggml_lightning_indexer to reduce peak memory usage.
  • CUDA backend support added for the lightning indexer with 4 heads, dispatching them to the vector kernel.
  • Metal backend updated to read head count from a function constant, allowing any head count to run without code changes.
  • Vulkan backend tiles the lightning indexer over keys and tokens using shared memory and fp16 dot products.

These changes improve memory efficiency for qwen4exp models with long contexts while maintaining compute buffer size and speed.