The llama.cpp project has released version b11390, which includes a fix for a CUDA MMQ memory fault that occurred when the number of experts was significantly larger than the micro-batch size.

  • Resolves a memory fault in the CUDA backend related to Mixture-of-Experts (MMQ) handling.
  • Provides binaries for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA 12/13, ROCm, OpenVINO, SYCL, Snapdragon), Windows (CPU, OpenCL, CUDA 12/13, Vulkan, OpenVINO, SYCL, ROCm), Android, and openEuler.
  • Includes an updated UI build.

This update ensures stability for users running Mixture-of-Experts models on CUDA hardware by preventing the specific memory access error.