Topic · Multimodal
lab NVIDIA Research · 12d ago · 8 views

NVIDIA introduces Spatial-IQ, a hierarchical diagnostic framework for multimodal model spatial reasoning

Researchers at NVIDIA have introduced Spatial-IQ, a diagnostic framework designed to deconstruct the spatial intelligence of multimodal large language models (MLLMs). Unlike existing benchmarks that treat models as black boxes, this approach decomposes object counting in stacked 3D structures into nine perceptual and cognitive sub-tasks aligned with human developmental stages.

media MarkTechPost · 4d ago · 1 view

NVIDIA releases Alpamayo 2 Super, a 34B open VLA model for autonomous driving

NVIDIA has released Alpamayo 2 Super, a 34-billion-parameter vision-language-action (VLA) model designed for robotaxis and autonomous driving under the OpenMDW-1.1 license. The model targets long-tail events by combining a 32B Cosmos 3 Super Reasoner backbone with a 2.3B diffusion-based action decoder to generate trajectories, causal explanations, and meta-actions from multi-camera video.