The llama.cpp project has released version b11205, which introduces a specific update to its CUDA backend. This release adds support for the Nemotron 3 Puzzle state size of 96 within the SSM scan implementation.

  • The update enables handling of Nemotron 3 models with a puzzle state size of 96 on CUDA hardware.
  • Binaries are provided for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, CUDA, ROCm, OpenVINO, SYCL, Snapdragon), Windows (CPU, CUDA, Vulkan, OpenVINO, SYCL, ROCm), Android, and iOS.
  • The release includes standard UI binaries alongside the platform-specific executables.

This update allows users to run Nemotron 3 models with specific state sizes on compatible NVIDIA GPUs using llama.cpp.