Multimodal
media MarkTechPost · 4d ago · 1 view

NVIDIA releases Alpamayo 2 Super, a 34B open VLA model for autonomous driving

NVIDIA has released Alpamayo 2 Super, a 34-billion-parameter vision-language-action (VLA) model designed for robotaxis and autonomous driving under the OpenMDW-1.1 license. The model targets long-tail events by combining a 32B Cosmos 3 Super Reasoner backbone with a 2.3B diffusion-based action decoder to generate trajectories, causal explanations, and meta-actions from multi-camera video.

lab NVIDIA Research · 12d ago · 8 views

NVIDIA introduces Spatial-IQ, a hierarchical diagnostic framework for multimodal model spatial reasoning

Researchers at NVIDIA have introduced Spatial-IQ, a diagnostic framework designed to deconstruct the spatial intelligence of multimodal large language models (MLLMs). Unlike existing benchmarks that treat models as black boxes, this approach decomposes object counting in stacked 3D structures into nine perceptual and cognitive sub-tasks aligned with human developmental stages.

arxiv arXiv cs.AI · 12d ago · 1 view

ClinFusion introduces vision-centric MLLM for holistic medical understanding

Researchers introduce ClinFusion, a vision-centric multimodal large language model designed to address the challenges of deploying AI in clinical practice by unifying 2D and 3D medical image understanding. The system features a compositional cascaded vision encoder with a Cascade Spatial-Aware Locality Fusion operator and includes a new evaluation framework comprising MedIF-Bench and region-of-interest-grounded metrics.

media MarkTechPost · 14d ago · 2 views

Induction Labs releases Photon-1, an imagination model trained on 18 years of video

Induction Labs has released Photon-1, a sparse 106B-A5B mixture-of-experts transformer that learns to simulate desktop environments and play games by predicting future frames from raw video without action labels. The model uses finite scalar quantization to compress each frame into 960 discrete tokens, achieving over 100 times better compression than existing representations while preserving text and layout details.