Researchers address temporal discontinuities in long-horizon video relighting by reframing the task as temporally conditioned latent domain translation. The framework enforces cross-chunk continuity by propagating target-domain latents across boundaries and utilizes masked target-domain self-conditioning to make this behavior learnable.
- Propagates target-domain latents across chunk boundaries to maintain temporal consistency.
- Employs masked target-domain self-conditioning, training the model to continue from temporally masked propagated context.
- Introduces warm-start prompting with a relit prompt anchor to establish the initial target-domain state.
- Provides a general interface for prompt-based relighting using a controllable generative model.
Experiments on in-the-wild long-horizon videos show markedly improved temporal consistency, with chunk-boundary artifacts largely reduced and unwanted appearance changes across chunks greatly suppressed.