- Black Forest Labs released FLUX 3, its first multimodal model generating video up to 20 seconds with synced audio.
- The same backbone powers FLUX-mimic, a robotics model already tested by Audi on its production line for flexible door seals.
- Only the open-weight “Dev” version is planned for later in 2026; video and action remain behind APIs and partner access for now.
- Human evaluators preferred FLUX 3 over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93%.
Black Forest Labs released FLUX 3 on Thursday, marking a shift from static image generation to video and audio. The German AI lab trained the model on images, video, and audio simultaneously inside one shared system, achieving multimodality.
The video side produces clips up to 20 seconds long with audio synced to on-screen events. In head‑to‑head evaluations, human reviewers preferred FLUX 3 over Runway Gen-4.5 77% of the time and over Luma Ray 3.2 93% of the time.
The company frames this as more than a content tool. “A model that only learns images can only generate images,” said co‑founder and CEO Robin Rombach.
That bet has materialized as FLUX-mimic, built with Zurich‑based mimic robotics. The system adds a lightweight decoder that translates video‑prediction into robot motions, and Audi is already testing it on tasks like fitting flexible door seals.
“Audi represents the kind of manufacturing partner we built FLUX-mimic for,” said mimic co‑founder Stephan‑Daniel Gravert. Audi’s Christoph Schneider added that robots now solve complex soft‑body manipulation work older machines could not handle.
BFL released FLUX.2 in November 2025 but it was less popular, while the original open‑source Flux was eventually surpassed by Alibaba’s Z-Image Turbo. FLUX 3 is BFL’s comeback, though video and action remain behind APIs and partner access, with image generation following in coming weeks.
✅ Follow BITNEWSBOT on Telegram, Facebook, LinkedIn, X.com, and Google News for instant updates.
