Researchers from NVIDIA, MIT, and the University of Oxford have introduced Physis-Lang, a framework that enhances video world models by integrating self-evolving physical language into captions. This approach treats physical language as an optimizable representation used for data curation, model training, and inference to correct physics errors in generated videos.

  • Physis-Lang adds a `physics_reasoning` field and a `physics_negative_prompt` to base captions, detailing entities, causes, and governing principles.
  • A critic-guided agent evolves the caption instructions while keeping the captioner frozen, validated on the new PhysCapBench benchmark of 246 videos.
  • The method uses language-guided data curation to retrieve physically relevant clips, resulting in a final training set of 183K videos.
  • Fine-tuning Cosmos3-Nano with Physis-Lang beats Google's Veo 3.1 on PhyGenBench (71.04 vs 65.63) and Physics-IQ Verified (43.41 vs 34.99).
  • The pipeline can be distilled into local Qwen3-VL-4B-Instruct models, reducing API costs from approximately $24.12K to $0 while maintaining significant physics gains.

The framework allows video models to explain why and how scenes unfold, improving physical consistency without requiring architectural changes or commercial APIs.