The llama.cpp project has released version b11390, which includes a fix for a CUDA MMQ memory fault that occurred when the number of experts was significantly larger than the micro-batch size.
- Resolves a memory fault in the CUDA backend related to Mixture-of-Experts (MMQ) handling.
- Provides binaries for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA 12/13, ROCm, OpenVINO, SYCL, Snapdragon), Windows (CPU, OpenCL, CUDA 12/13, Vulkan, OpenVINO, SYCL, ROCm), Android, and openEuler.
- Includes an updated UI build.
This update ensures stability for users running Mixture-of-Experts models on CUDA hardware by preventing the specific memory access error.