> ## Content Index
> Fetch the complete content index at: https://www.metatalks.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Black Forest Labs' first video model FLUX 3 doubles as a robot controller on Audi's line
- URL: https://www.metatalks.ai/black-forest-labs-first-video-model-flux-3-doubles-as-a-robot-controller-on-audis-line/
- Published: 2026-07-24T13:11:57.000Z
- Updated: 2026-07-24T13:11:57.000Z
- Author: Al
- Tags: News, Physical AI, Frontier Models, #newswire

**The same backbone behind FLUX 3's video now runs FLUX-mimic, a robot controller the lab says learns a new factory job from a fraction of the demonstration footage usually required.**

Black Forest Labs granted a small group of testers early access this week to FLUX 3, [its first model that generates video rather than only stills](https://bfl.ai/blog/flux-3?ref=metatalks.ai). Trained on images, video and audio together in one system, it turns out clips of up to 20 seconds that arrive with a soundtrack locked to the action — spoken dialogue, effect noises and ambient background sound.

The same FLUX 3 backbone also drives FLUX-mimic, a robot-control model built with Zurich-based mimic robotics. A preliminary build is already operating on robots on Audi's assembly line, with mimic robotics among the earliest partners handed advance access to the model.

FLUX-mimic bolts a compact decoder onto the part of FLUX 3 that forecasts video, turning what the model has internally worked out about how objects move into the real motions a robot carries out. The bet, as co-founder and chief executive Robin Rombach frames it, is that a model taught only on images can do nothing but make images, whereas one taught to anticipate video also absorbs the physics beneath it — weight, contact and timing — precisely what a machine needs to get around in the physical world.

On Audi's line, according to the carmaker's Christoph Schneider, the robots can now take on intricate handling of soft, pliable materials — work earlier machines were unable to do. Black Forest Labs says the system picks up a fresh task from [roughly 30 minutes of demonstration footage, against the more than 30 hours usually required](https://bfl.ai/blog/flux-3-mimic%29with?ref=metatalks.ai), and reacts in about 101 milliseconds.

In the lab's own preference tests, evaluators favored FLUX 3's footage over rivals — 77% of the time against Runway Gen-4.5, 93% against Luma Ray 3.2 and 52% against Gemini Omni and Seedance. Those figures come from side-by-side comparisons where viewers simply pick the clip that looks and sounds more convincing, not from a fixed scoring rubric.

For now the video and action components are open only to a limited group through APIs and select partners; image generation is expected in the coming weeks, and the open-weight FLUX 3 Dev, built compact enough to run on the equipment found in factories, isn't foreseen until later.