The llama.cpp project has released build b10344, which introduces Multi-Token Prediction (MTP) support specifically for the Nemotron and Nemotron Nano models.
- Adds MTP support for Nemotron Nano model.
- Introduces mtp_flags configuration for the Nemotron model.
- Provides binaries for macOS (Apple Silicon and Intel), iOS, Linux (CPU, Vulkan, ROCm, OpenVINO, SYCL), Android, Windows (CPU, CUDA 12/13, Vulkan, OpenCL, OpenVINO, SYCL, HIP), and openEuler.
This update enables users to leverage MTP capabilities with Nemotron models across a wide range of hardware platforms and operating systems.