The llama.cpp project has released build b11025, which includes an extension of Nemotron Multi-Token Prediction (MTP) support.

  • The update addresses the first fix for this feature and removes unnecessary declarations.
  • Binaries are provided for macOS (Apple Silicon and Intel), iOS, Linux (CPU, Vulkan, CUDA 12/13, ROCm, OpenVINO, SYCL), Android, Windows (CPU, OpenCL, CUDA 12/13, Vulkan, OpenVINO, SYCL, ROCm), and openEuler.
  • A standalone UI build is also available for download.