The llama.cpp project has released version b11205, which introduces a specific update to its CUDA backend. This release adds support for the Nemotron 3 Puzzle state size of 96 within the SSM scan implementation.
- The update enables handling of Nemotron 3 models with a puzzle state size of 96 on CUDA hardware.
- Binaries are provided for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA, ROCm, OpenVINO, SYCL, Snapdragon), Windows (CPU, CUDA, Vulkan, OpenVINO, SYCL, ROCm), Android, and iOS.
- The release includes standard UI binaries alongside the platform-specific executables.
This update allows users to run Nemotron 3 models with specific state sizes on compatible NVIDIA GPUs using llama.cpp.