Shanghai AI Laboratory has introduced InternW0, the first instantiation of its InternW physical world model series, designed to maintain actionable predictions in dynamic environments. The model jointly learns future visual dynamics and continuous robot control using an asymmetric video-action architecture with flow matching.

  • InternW0 employs a high-capacity video expert for long-horizon context and a lightweight action expert for faster timescales.
  • It reuses layerwise K/V states and adapts them to new observations via context routing, avoiding regeneration for every update.
  • Training utilized approximately 7,200 hours of heterogeneous robot and egocentric data, including the 275-hour EgoLab dataset.
  • Evaluation included a 15-stage metal-organic framework synthesis workflow and 5-stage contact-aware dexterous manipulation.

These results advance scalable, asynchronous, and science-native physical world models for universal and efficient real-world interactions.