Black Forest Labs has released FLUX 3, a multimodal foundation model that learns from images, videos, and audio within a single architecture. It is the first FLUX model to ship video, audio, and action prediction from one set of weights.

  • The model builds on the Self-Flow method, combining flow matching with self-supervised feature reconstruction.
  • FLUX 3 Video generates clips up to 20 seconds long in a single generation with native audio.
  • Supported modes include text-to-video, image-to-video, video-to-video, keyframe-to-video, and generative video-audio continuation.
  • In human preference tests for 10-second 720p clips, FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and Runway Gen-4.5 in 77%.

The release aims to provide a unified model where modalities constrain each other, ensuring sound matches impact and motion obeys mass.