NVIDIA has released two open-weight models, d1-3B and d1-omni-600M, as part of its new d1 decision model family. Unlike generative models that produce tokens, these models generate answers in a single forward pass, enabling extremely low-latency inference across NVIDIA's full stack from data centers to edge devices.
- d1-3B scores 48.57 on the Decision Index v0.2.1 public split, outperforming every model under 10B and matching Decider 35B-A3B at a fraction of the size.
- d1-omni-600M is an experimental checkpoint handling text, image, and audio inputs, scoring 15.95 on the same index.
- d1-3B achieves inference latencies of 8 ms on NVIDIA GeForce RTX 4090, 16 ms on Jetson AGX Thor, and 26 ms on Jetson AGX Orin.
- Both models support day-one compatibility with llama.cpp and run on Apple M5 Pro, AMD MI325X, and various NVIDIA Jetson devices.
These open-weight models allow users to download, fine-tune, and deploy fast, structured decision-making capabilities for multimodal inputs without restrictions.