The llama.cpp b11265 release includes a fix for the Vulkan backend's handling of Mixture of Experts (MoE) models. The update corrects how `mat_mul_id` selects matrix multiplication tiles, addressing a performance issue where workers were left idle during dispatch.

  • Previously, tile selection relied on total token count, which was incorrect for MoE dispatch grids where the true N per workgroup is per-expert rows.
  • On models like Sarvam 30B with pp128, this error caused the picker to select tiles for approximately 6 live rows instead of 128.
  • The fix ensures that all workers in a group have work to do, eliminating wasted time that previously accounted for 55% of the job's duration.

This change improves inference efficiency for MoE models on Vulkan by ensuring proper workload distribution across GPU workers.