A new method evaluates physical consistency in generated videos without requiring human voting or ground-truth references. It uses DROID-SLAM and SEA-RAFT to detect inconsistencies, improving task success rates by over 8% and enabling spatio-temporal localization of physical artifacts.
Reference-Free Assessment of Physical Consistency in Video Generation
Testing Seedance 2.5 character consistency across multiple fashion looks
A user tested the Seedance 2.5 video generation model using a short fashion and editorial sequence featuring a single character in various outfits, compositions, and visual styles.
CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning
Researchers propose CineCap, a framework that combines structured reasoning with spatio-temporal anchors and reinforcement learning to improve cinematographic video captioning. The method grounds professional film-language descriptions in explicit visual evidence while balancing descriptive completeness and factual correctness.
Reference-Free Assessment of Physical Consistency in World Model-based Video Generation
The authors introduce reference-free measures for evaluating the physical consistency of generated videos by combining relative and absolute fidelity assessments. This approach addresses the gap in physical fidelity that often prevents video generation tools like WorldGym or WorldEval from accurately reproducing real-world task success rates for VLA models. Unlike existing methods requiring costly human voting or unavailable ground-truth references, the new framework utilizes DROID-SLAM and SEA-RAFT to quantify inconsistencies. Motivated by WorldScore, the relative consistency assessment filters videos to improve task success rates by over 8%. Additionally, the absolute assessment enables spatio-temporal localization to visualize when and where physical artifacts occur in the generated content.
Nvidia acquires Hugging Face; OpenAI releases GPT-6 Astra; Microsoft launches MAI-Transcribe-2
In early September 2026, Nvidia confirmed its $12.93 billion acquisition of Hugging Face, while OpenAI released GPT-6 Astra and Microsoft launched the MAI-Transcribe-2 speech recognition model.
Anthropic releases Claude Fable 5.1 and Mythos 5.1 system card
Anthropic has published the system card for Claude Fable 5.1 and Mythos 5.1, detailing their safety profiles, alignment risks, and capabilities relative to previous models.