The llama.cpp project has released build b10344, which introduces Multi-Token Prediction (MTP) support specifically for the Nemotron and Nemotron Nano models.

  • Adds MTP support for Nemotron Nano model.
  • Introduces mtp_flags configuration for the Nemotron model.
  • Provides binaries for macOS (Apple Silicon and Intel), iOS, Linux (CPU, Vulkan, ROCm, OpenVINO, SYCL), Android, Windows (CPU, CUDA 12/13, Vulkan, OpenCL, OpenVINO, SYCL, HIP), and openEuler.

This update enables users to leverage MTP capabilities with Nemotron models across a wide range of hardware platforms and operating systems.