FLUX 3
Multimodal Flow Models as Visual Intelligence
FLUX 3 jointly learns from images, videos, and audio within a unified architecture. It perceives physical dynamics, captures natural acoustic phenomena, and powers content creation and embodied AI.
* Note: This website is an independent showcase page and is not affiliated with Black Forest Labs.

Perceive, Predict & Act
No single modality provides a complete description. Images capture spatial structure, videos add temporal dynamics, audio reveals acoustics, language adds instruction.
Self-Flow: The Multimodal Foundation
Rather than training isolated diffusion models for image or text, FLUX 3 builds on Self-Flow to jointly learn from images, video, and audio as mutual physical constraints of one underlying reality.
Spatial Structure
Images capture crisp spatial geometry, object relationships, and lighting balance at a precise frame of time.
Temporal Dynamics
Videos restore the temporal axis, establishing physical law continuity, object inertia, and motion flow.
Acoustic Causality
Audio enforces physical contact acoustics, mechanical impact timing, and environmental resonance.
Action & Embodiment
Action vectors allow the backbone to predict physical robot trajectories in industrial production lines.
Unified Multimodal Flow Matching Model

Self-Flow vs. Standard Flow Matching (FM)
Empirical evaluation comparing Self-Flow against traditional Flow Matching.

One Model. Multiple Frontier Capabilities.
FLUX 3 generates across vision, acoustics, and action in a single inference call.
FLUX 3 Video with Native Spatial Audio
FLUX 3 generates long-form clips up to 20 seconds in a single pass, featuring tightly aligned multi-channel soundscapes, facial expressions, and complex spatial physics.
Human Preference Win Rates
Blind head-to-head human preference evaluations generated over 10-second text-to-video clips in 720p with native audio.
FLUX 3 vs. Leading Video Foundation Models

Early Access Phase Note: As FLUX 3 and its harness are actively undergoing fine-tuning, these early evaluations reflect midtraining checkpoints. Further benchmark gains are expected as models transition from early access to general rollout.
Try FLUX 3 Generation Pipeline
Simulate text, vision, and acoustic flow matching prompts in real-time.
FLUX 3 Capabilities Launch Roadmap
Gradually releasing multimodal capabilities following rigorous safety testing and feedback loops.
FLUX 3 Video + Audio
Full API and private weight access for 20s native video generation with spatial audio sync.
FLUX 3 Action & FLUX-mimic
Robotic manipulation trajectory prediction co-developed with mimic robotics and Audi.
FLUX 3 Image
Enhanced text-to-image synthesis, multi-language typography accuracy, and non-destructive editing.
FLUX 3 Dev (Open Weights)
Open-weight multimodal foundation model backbone for community researchers and developers.
FLUX 3 Multimodal FAQ
Answers to common questions regarding FLUX 3, Self-Flow architecture, and access roadmap.
FLUX 3 is a multimodal frontier model that jointly learns from images, video, and audio within a unified flow architecture. Unlike single-modality models, FLUX 3 treats spatial structure, temporal motion, and acoustic physics as mutual constraints of one underlying reality.
Self-Flow is Black Forest Labs' approach for efficiently aligning multimodal generation and understanding within the same flow matching backbone. It allows joint training across video, images, audio, and physical action vectors while achieving significantly lower generation errors.
FLUX 3 supports native 20-second video generation with multi-channel synchronized spatial audio, high-precision multilingual text rendering on images, video-to-video style transfer, keyframe transitions, and robotic action trajectory prediction.
Yes, as part of the launch roadmap, FLUX 3 Dev will be made available as an open-weight multimodal backbone for content creation and action prediction research.
Developed in partnership with mimic robotics and tested on Audi production lines, FLUX-mimic uses FLUX 3's pretrained video dynamics backbone as a physical foundation, allowing action models to be fine-tuned for dexterous robot manipulation with limited data.