2026-07-23

FLUX 3: Black Forest Labs Bets Image, Video, and Robots Are One Problem

AIRobotics🌍 Europe

Black Forest Labs β€” the ~100-person Freiburg lab founded by the original Stable Diffusion authors, now valued at $3.25B after its December Series B β€” announced FLUX 3 on July 23, 2026. And the interesting thing isn't the image quality. It's the thesis.

From image model to "visual intelligence"

FLUX.1 and FLUX.2 were image models, full stop β€” excellent ones, open-weight, the backbone of half the creative-tool ecosystem. FLUX 3 is framed differently: a single unified architecture jointly trained on images, video, and audio, which "can also be extended to predict actions." BFL's claim, in CEO Robin Rombach's words, is that true intelligence means "perceiving the world: predicting how it will change, taking action, and learning from the results" β€” and that image generation, video generation, world models, and robotics are all "expressions of the same underlying capability."

The concrete version of that claim: BFL says generative video and action prediction don't require separate foundations β€” the same model extends to predicting robot actions without sacrificing its video capabilities, and each training modality strengthens the others.

Three variants, staggered rollout

  • FLUX 3 Video β€” video generation with optional native synchronized audio, in early access now. BFL says it "already leads in early evaluations against frontier video models," with particular strength in human facial expressions, associating sounds with physical events, and multilingual capability β€” though no benchmark numbers or arena rankings have been published yet; those are promised "alongside broader availability."
  • FLUX 3 Action β€” the robotics arm of the thesis, also in early access. The flagship deployment is FLUX-mimic, a video-action model built with Swiss startup mimic robotics, already in testing with manufacturers including Audi β€” dual-arm dexterous manipulators doing soft-body manipulation on production lines. That's not a demo reel; that's a car factory.
  • FLUX 3 Image β€” precise editing with product and material consistency, rolling out in the coming weeks.

Parameter counts, licenses, and pricing are all undisclosed for now, but BFL has committed to faster and open-weight versions of FLUX 3 later this year β€” with an explicitly robotics-flavored rationale: local, low-latency deployment for robot control systems, and fine-tuning on customer data. Early testers include Canva, Burda, Freepik's Magnific, Krea, and Picsart.

Why this is worth watching

Plenty of labs do video generation, and plenty do robot policies β€” the bet that they're the same model is rarer, and BFL is unusually well-positioned to test it: a European lab with genuine open-weight credibility, frontier image quality as the starting point, and a real industrial partner putting the action variant on actual production lines. It also rhymes with where the rest of the field has been drifting this year β€” NVIDIA's Cosmos 3 unified predict/transfer/reason/act into one architecture for physical AI just last week. If "visual intelligence" pans out as a single foundation, the image-model labs may turn out to have been robotics labs all along.

Links

FLUX 3: Black Forest Labs Bets Image, Video, and Robots Are One Problem | Laura Martel