Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction

TL;DR AI
2 min readKey summary
Black Forest Labs launched FLUX 3, a multimodal foundation model built on Self-Flow.
The model is trained jointly on images, video, and audio, and includes early-access video and action capabilities.
A related policy model, FLUX-mimic, suggests the same backbone could also support robot action prediction.
The release highlights a single shared architecture for generative media and robotics, with potential gains in cross-modal learning.



