NVIDIA has released Alpamayo 2 Super, a 34-billion-parameter vision-language-action (VLA) model designed for robotaxis and autonomous driving under the OpenMDW-1.1 license. The model targets long-tail events by combining a 32B Cosmos 3 Super Reasoner backbone with a 2.3B diffusion-based action decoder to generate trajectories, causal explanations, and meta-actions from multi-camera video.
- Alpamayo 2 Super ranks first on LingoQA with a Lingo-Judge score of 79.2, outperforming Qwen2.5-VL 72B, Gemini 2.5 Pro, and GPT-4o.
- The model was trained on approximately 115,000 hours of driving video and over one billion images, including 3.7 million Chain-of-Causation traces.
- It produces five outputs from a single pass: trajectory, CoC trace, meta-action, reasoning auto-labels, and visual question answering with 2D grounding.
- Closed-loop evaluation on AlpaSim yielded a score of 1.50 ± 0.13, while open-loop testing achieved a minADE₆ of 0.911m at 6.4 seconds.
The release allows commercial use and redistribution without additional permission, and NVIDIA states the model can compress annotation cycles from months to days when used as an autolabeler for proprietary fleet data.