Dyna Robotics has released Dyna-2, a world-action model for robot manipulation that was pre-trained on more than one million hours of egocentric human video. The company demonstrates that ordinary human video can substitute for action-labelled data by establishing a scaling law from 1,000 to 1,000,000 hours.
- Dyna-2 is a generative model using a video-diffusion backbone that jointly denoises future video and action chunks.
- The model exhibits a scaling law on human data that transfers zero-shot to unseen robot tasks across two stationary bimanual platforms.
- Joint denoising of video and action outperforms action-only training, with video specifically aiding cross-embodiment generalization.
- Post-training on 14 tasks showed mean normalized scores rising from 20% to 53% as data scaled to one million hours.
- Dyna-2 passed production criteria at 87% of unseen customer sites compared to 46% for the previous Dyna-1 model.
The research indicates that video prediction drives transfer to robot data, allowing the system to learn from vast amounts of unlabeled human experience rather than relying on costly teleoperation.