Black Forest Labs, the German artificial intelligence powerhouse renowned for its groundbreaking FLUX line of image generators, marked a significant milestone on Thursday with the release of FLUX 3. This latest iteration of the company’s flagship model introduces a transformative capability: native video generation, moving beyond the still images that defined its predecessors. In a departure from conventional AI development, FLUX 3 was meticulously trained on a unified dataset encompassing images, video, and audio simultaneously, all within a single, cohesive system. This integrated approach represents a powerful leap into true multimodality, where an AI model learns diverse types of information in concert, fostering a deeper, more holistic understanding of the world it seeks to represent and interact with.
The Dawn of Multimodal Video Generation
At the core of FLUX 3’s revolutionary offering is its advanced video generation capability, which stands as the headline feature of this release. The model is capable of producing video clips up to 20 seconds in duration, a notable length in the rapidly evolving landscape of generative AI video. What truly distinguishes FLUX 3, however, is its integrated audio generation, meticulously synced with the visual content. This includes contextually appropriate dialogue, realistic sound effects, and immersive ambient noise, all designed to enhance the coherence and realism of the generated footage. This synchronized audiovisual output addresses a critical challenge in AI-generated video, where a disconnect between sight and sound can often break immersion and believability.
Early evaluations underscore FLUX 3’s formidable performance against established industry benchmarks. In head-to-head comparisons conducted by human reviewers, FLUX 3’s output was preferred over Runway Gen-4.5 in a remarkable 77% of instances. Its superiority was even more pronounced against Luma Ray 3.2, with FLUX 3 winning 93% of comparisons. While the margin was narrower, FLUX 3 still demonstrated a slight edge over prominent models like Gemini Omni and Seedance, besting them in 52% of evaluations. It is crucial to note that these figures represent preference tests, where evaluators simply selected the clip that appeared and sounded more convincing. This qualitative assessment, though not a fixed scoring rubric, provides valuable insight into the perceived quality and realism of FLUX 3’s generated content. Beyond its pioneering video capabilities, FLUX 3 also maintains, and arguably enhances, its legacy in still image generation. Black Forest Labs showcased a diverse array of images, demonstrating FLUX 3’s remarkable versatility across various artistic styles, extending far beyond mere photorealism. This dual proficiency positions FLUX 3 as a comprehensive creative tool, capable of meeting a broad spectrum of visual content demands.
Beyond Pixels: Understanding the Physics of Reality
The strategic decision to train FLUX 3 on multimodal data – images, video, and audio together – is not merely about enhancing content creation. According to Robin Rombach, co-founder and CEO of Black Forest Labs, this approach serves a much grander ambition. "A model that only learns images can only generate images," Rombach stated, emphasizing the limitations of unimodal systems. The company’s profound belief is that by learning to predict video, an AI system simultaneously acquires an intrinsic understanding of the underlying physics governing the visual world. This encompasses fundamental concepts such as weight, contact dynamics, and precise timing – elements essential for comprehending how objects interact and move in physical space.
This deeper comprehension of physical laws is precisely what a machine requires to navigate and interact effectively within the real world. It marks a paradigm shift from AI as purely a digital content creator to AI as an intelligent agent capable of physical action. This foundational bet by Black Forest Labs has a concrete manifestation in FLUX-mimic, an innovative collaboration with Zurich-based mimic robotics. FLUX-mimic ingeniously integrates FLUX 3’s advanced video-prediction engine with a lightweight "decoder." This small yet crucial add-on component serves as the bridge, translating the model’s internal, abstract understanding of motion and physics into tangible, real-world robot movements.

The practical implications of FLUX-mimic are already being explored in critical industrial applications. Automotive giant Audi is actively testing the system for tasks traditionally deemed too complex for conventional automation, such as fitting flexible door seals onto vehicles. These tasks involve delicate manipulation of pliable materials, requiring a nuanced understanding of force, shape, and elasticity – precisely the kind of "soft-body manipulation" that older, rigid robotic systems have consistently struggled to handle. Stephan-Daniel Gravert, co-founder of mimic robotics, underscored this synergy, stating, "Audi represents the kind of manufacturing partner we built FLUX-mimic for." Christoph Schneider of Audi further validated the technology’s impact, confirming that the robots powered by FLUX-mimic now "solve complex soft-body manipulation work" that was previously beyond the scope of existing machinery. The entire FLUX-mimic system boasts an impressive reaction time of approximately 101 milliseconds, a responsiveness that places it in the vicinity of human visual reflexes, paving the way for highly adaptive and precise robotic operations.
Black Forest Labs: A Chronicle of Disruption and Innovation
The ascendancy of the FLUX models and Black Forest Labs itself is a story of rapid innovation and strategic disruption within the burgeoning AI landscape. The company was founded in August 2024 by a cohort of veteran researchers who played pivotal roles in developing the original Stable Diffusion models at Stability AI. Their departure from Stability AI signaled an intent to push the boundaries of generative AI in new directions, and they wasted no time in making their mark.
Upon their initial release, the early FLUX models quickly garnered attention by demonstrably outperforming established players like MidJourney and even Stability AI’s own much-anticipated but ultimately underwhelming Stable Diffusion 3. This immediate success established Black Forest Labs as a formidable challenger. Their open-source offerings, Flux Dev and Schnell, were particularly impactful, rapidly seizing the coveted title of "best open-source image generator." This achievement was especially significant as the AI art community had widely expected Stable Diffusion 3.5, Stability AI’s subsequent attempt to rectify SD3’s shortcomings, to claim that mantle. However, it never did.
The momentum continued with the commercial release of FLUX 1.1 Pro in October of the same year. This proprietary model further solidified BFL’s reputation, topping the prestigious Artificial Analysis image arena, a respected benchmark for AI image generation quality. This marked a clear differentiation between BFL’s open-source contributions and its commercially focused, high-performance models.
In November 2025, Black Forest Labs released FLUX 2. While still an advanced model, it did not achieve the same level of widespread popularity or critical acclaim as its predecessor. The open-source crown, initially held by the original Flux, eventually passed to Alibaba’s Z-Image Turbo in late 2025. Z-Image Turbo managed to match Flux’s quality while operating efficiently on lower-end consumer graphics cards, a crucial factor for broader accessibility and adoption among individual AI artists and developers. A user on CivitAI, a prominent platform for AI art models, famously remarked at the time, "This is what SD3 was supposed to be," reflecting the community’s desire for high-quality, accessible open-source models.
FLUX 3, therefore, represents not just an evolutionary step but a powerful comeback for Black Forest Labs, reaffirming its position at the vanguard of generative AI innovation. Its multimodal architecture and groundbreaking video capabilities are designed to reclaim leadership and redefine expectations within the industry.

Strategic Rollout and Future Implications
The release of FLUX 3 is being managed through a phased and strategic rollout. Currently, its advanced Video and Action functionalities are available in early access via APIs and to select partners, with mimic robotics being a prime example of such a collaborator. The highly anticipated image generation capabilities of FLUX 3 are slated to become broadly available "in the coming weeks," allowing a wider audience of content creators and developers to leverage its enhanced visual output. For the broader AI community and individual users, Black Forest Labs plans to release an open-weight Dev version later in 2026. This tier will be the only one specifically designed for local deployment and use, catering to developers and enthusiasts who wish to integrate and experiment with FLUX 3’s core technology on their own hardware.
The implications of FLUX 3 and its underlying multimodal philosophy are far-reaching, poised to reshape several industries and accelerate the trajectory of artificial intelligence. In the creative sector, the ability to generate high-quality, synchronized video and audio clips up to 20 seconds long could revolutionize content creation for film, advertising, gaming, and digital media. It promises to democratize video production, enabling creators with limited resources to produce sophisticated visual narratives with unprecedented efficiency. The advancements in generating coherent, physically realistic scenes could lead to more believable virtual worlds and characters, significantly impacting the entertainment and metaverse industries.
Concurrently, the integration of FLUX 3’s capabilities with robotics through FLUX-mimic heralds a new era for industrial automation. The ability for robots to "learn the physics underneath" through multimodal data ingestion signifies a critical step towards more intelligent, adaptable, and versatile robotic systems. Tasks that are currently challenging or impossible for traditional automation – those requiring fine motor skills, adaptability to variable conditions, and an intuitive understanding of material properties – could become routine. This has profound implications for manufacturing, logistics, healthcare, and hazardous environments, promising increased efficiency, safety, and operational flexibility.
The competitive landscape of AI development is also undergoing a significant shift. FLUX 3’s success in multimodal integration sets a new benchmark, compelling other major players to accelerate their own research and development in this area. Multimodality is emerging as the next frontier in the race towards more general and human-like artificial intelligence. As models become increasingly adept at processing and synthesizing information across various modalities, they move closer to achieving a comprehensive understanding of human communication and the physical world. While the ethical considerations surrounding generative AI, such as potential misuse in deepfakes or the broader impact on employment, remain important subjects of ongoing discussion, the immediate focus for Black Forest Labs is on empowering creators and industries with tools that enhance productivity and unlock new possibilities. FLUX 3 is not just an upgrade; it is a declaration of intent, positioning Black Forest Labs at the forefront of AI’s evolution from specialized tools to truly intelligent, multifaceted systems.















