The same backbone behind FLUX 3's video now runs FLUX-mimic, a robot controller the lab says learns a new factory job from a fraction of the demonstration footage usually required.

Black Forest Labs granted a small group of testers early access this week to FLUX 3, its first model that generates video rather than only stills. Trained on images, video and audio together in one system, it turns out clips of up to 20 seconds that arrive with a soundtrack locked to the action — spoken dialogue, effect noises and ambient background sound.

The same FLUX 3 backbone also drives FLUX-mimic, a robot-control model built with Zurich-based mimic robotics. A preliminary build is already operating on robots on Audi's assembly line, with mimic robotics among the earliest partners handed advance access to the model.

FLUX-mimic bolts a compact decoder onto the part of FLUX 3 that forecasts video, turning what the model has internally worked out about how objects move into the real motions a robot carries out. The bet, as co-founder and chief executive Robin Rombach frames it, is that a model taught only on images can do nothing but make images, whereas one taught to anticipate video also absorbs the physics beneath it — weight, contact and timing — precisely what a machine needs to get around in the physical world.

On Audi's line, according to the carmaker's Christoph Schneider, the robots can now take on intricate handling of soft, pliable materials — work earlier machines were unable to do. Black Forest Labs says the system picks up a fresh task from roughly 30 minutes of demonstration footage, against the more than 30 hours usually required, and reacts in about 101 milliseconds.

In the lab's own preference tests, evaluators favored FLUX 3's footage over rivals — 77% of the time against Runway Gen-4.5, 93% against Luma Ray 3.2 and 52% against Gemini Omni and Seedance. Those figures come from side-by-side comparisons where viewers simply pick the clip that looks and sounds more convincing, not from a fixed scoring rubric.

For now the video and action components are open only to a limited group through APIs and select partners; image generation is expected in the coming weeks, and the open-weight FLUX 3 Dev, built compact enough to run on the equipment found in factories, isn't foreseen until later.