Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction

TL;DR AI
2 min readKey summary
Black Forest Labs has launched FLUX 3, a multimodal model that handles images, video, audio, and robot actions from a single backbone.
Built on the company’s Self-Flow approach, FLUX 3 is jointly trained to improve consistency across what is seen, heard, and predicted.
FLUX 3 Video supports clips up to 20 seconds with native audio, while video and action capabilities are rolling out in early access.
The release points to broader use in generative media and embodied AI, where cross-modal physical consistency matters.
