Multimodal AI image, video, and audio generator with world understanding.
|
Founded year:
|
2025
|
|
Country:
|
United States of America
|
|
Funding rounds:
|
Not set
|
|
Total funding amount:
|
Not set
|
Description
Flux 3 is a multimodal foundation model that learns jointly from images, video, and audio to understand real-world dynamics. It generates diverse videos up to 20 seconds with native audio, offers advanced image synthesis and editing, and extends into action prediction for physical AI applications. The model excels in human facial expressions, sound matching, and multilingual generation, with a staged release plan for APIs and open-weight access.