Researchers address temporal discontinuities in long-horizon video relighting by reframing the task as temporally conditioned latent domain translation. The framework enforces cross-chunk continuity by propagating target-domain latents across boundaries and utilizes masked target-domain self-conditioning to make this behavior learnable.

  • Propagates target-domain latents across chunk boundaries to maintain temporal consistency.
  • Employs masked target-domain self-conditioning, training the model to continue from temporally masked propagated context.
  • Introduces warm-start prompting with a relit prompt anchor to establish the initial target-domain state.
  • Provides a general interface for prompt-based relighting using a controllable generative model.

Experiments on in-the-wild long-horizon videos show markedly improved temporal consistency, with chunk-boundary artifacts largely reduced and unwanted appearance changes across chunks greatly suppressed.