The llama.cpp project has released build b11370, which includes a CUDA optimization to fuse shared experts into the Mixture of MoE with Vector Quantization (MMVQ) architecture. This change is part of pull request #29184 and also addresses buffer null checks and argument handling for stride_col_dst.

  • Fuses shared experts into MMVQ for CUDA backends.
  • Adds null checks for buffers to improve stability.
  • Moves stride_col_dst to fusion arguments.

This release provides updated binaries for macOS, Linux, Windows, Android, and iOS across various hardware accelerators including CPU, Vulkan, ROCm, OpenVINO, SYCL, and Snapdragon.