Robotics & Physical AI
mimic robotics Introduces FLUX-mimic to Bring Video-Action Models to Audi’s Factory Floor

mimic robotics has introduced FLUX-mimic, a new Video-Action Model designed to help industrial robots learn complex manipulation tasks from substantially smaller amounts of demonstration data.
Developed in collaboration with visual AI company Black Forest Labs, the system is already being tested and deployed with Audi across manufacturing tasks that have historically been difficult to automate, including work involving flexible materials, changing part configurations, and precise manipulation.
The announcement points to a potential shift in physical AI. Instead of relying primarily on large collections of robot demonstrations, FLUX-mimic attempts to transfer knowledge already learned by a generative video model into a system capable of controlling physical machines.
A New Approach to Robot Learning
Many of today’s general-purpose robotics systems are based on Vision-Language-Action models, commonly known as VLAs. These systems typically combine a model trained on images and language with an additional component that predicts the movements a robot should make.
This approach gives robots a degree of semantic understanding, such as recognizing an object or interpreting a written instruction. However, static images do not fully capture how objects bend, move, collide, or respond to physical contact.
The missing knowledge must therefore be learned from robot demonstrations, which can be expensive and time-consuming to collect. Each demonstration requires a physical robot, an operator, a suitable workspace, and carefully recorded action data.
FLUX-mimic takes a different route. It uses the visual representations produced by Black Forest Labs’ FLUX 3 model as the foundation for predicting how a task should unfold over time. Black Forest Labs describes FLUX 3 as a multimodal model capable of generating images, video, audio, and action, rather than a system limited to static image generation.
An action decoder then translates the model’s visual prediction into commands that a robot can execute. In practical terms, the system first estimates what successfully completing a task should look like and then determines which physical movements are needed to produce that outcome.
Building on the mimic-video Architecture
FLUX-mimic builds on mimic-video, the company’s earlier research into using pretrained video models as the backbone for robotic control.
Rather than asking a conventional vision-language model to infer motion from individual frames, mimic-video creates a latent visual plan representing how the task may develop. A separate action decoder uses that internal representation to generate robot movements.
In its research paper, the company reported that mimic-video achieved approximately ten times greater sample efficiency and twice the training convergence speed of a comparable VLA architecture across its evaluations. The real-world system was trained using roughly 500 trajectories from a dexterous two-handed robot setup.
FLUX-mimic extends that concept using the architecture behind FLUX 3 and a design created specifically for physical AI.
According to mimic robotics, tasks that might previously have required 30 hours or more of robot demonstrations can reach production-ready performance with as little as 30 minutes of data. That figure represents the company’s reported results and will ultimately need to be evaluated across a wider range of factories, hardware configurations, and production conditions.
Even so, reducing the data requirement by that degree could materially change the economics of industrial robotics. Deployment costs frequently extend far beyond the robot itself, with considerable resources devoted to programming, integration, simulation, testing, and re-engineering whenever a product or process changes.
Audi Provides a Demanding Industrial Test
Audi is working with mimic robotics to evaluate FLUX-mimic in real manufacturing environments, where automation systems must contend with product variation, tight production requirements, and tasks that cannot always be reduced to repetitive movements.
The automaker says the robots have performed complex soft-body manipulation work that would have been extremely difficult to address with conventional robotics. Soft materials are especially challenging because their shape changes during handling, making it difficult to define every movement through fixed coordinates or rules.
Audi’s broader automation research illustrates why adaptability is becoming increasingly important. At its Böllinger Höfe facility, the company has been testing artificial intelligence, mobile robots, computer vision, and new picking technologies in a real-world laboratory. Audi has noted that the high level of customization in vehicles such as the e-tron GT creates complex logistics processes involving large numbers of different parts.
Traditional industrial robots remain highly effective when an environment is structured and every object arrives in a predictable position. They become less economical when parts, materials, or product variants change frequently enough to require repeated engineering work.
Learning-based systems could make those lower-volume and higher-variation processes more practical to automate. Instead of writing a new program for every task, manufacturers could demonstrate the desired behavior and allow the system to learn how to reproduce it.
Combining Models, Hardware, and Human Demonstrations
The model is part of a broader full-stack robotics platform being developed by mimic robotics.
The company builds its own dexterous robotic hand hardware, wearable data-collection equipment, robot-learning models, and deployment infrastructure. Its wearable system is designed to record human hand movements in a format that can be transferred more directly to the company’s robotic hands.
This focus on hands rather than complete humanoid robots reflects a more targeted approach to industrial automation. Dexterous hands can be attached to conventional robotic arms already used in factories, reducing the need to introduce an entirely new robotic form factor. ETH Zurich describes the approach as combining human-like manipulation with established industrial hardware to address tasks that require adaptability and self-correction.
Founded as an ETH Zurich spin-off in 2024, mimic robotics has grown to more than 50 employees across Zurich and San Francisco. Its team spans model research, robotics, hardware development, data operations, and industrial deployment.
The Wider Implications for Factory Automation
The significance of FLUX-mimic will depend less on isolated demonstrations than on whether the technology can maintain high success rates over thousands of production cycles.
Factories require consistent performance, predictable cycle times, safe operation around employees, rapid recovery from errors, and integration with existing manufacturing systems. A model that learns quickly but behaves unpredictably would offer little advantage over conventional automation.
However, if video-action models can retain their reported data efficiency while meeting industrial reliability requirements, they could expand robotics into categories of work that have remained manual because they were too variable or costly to program.
That could include assembling flexible components, handling mixed inventory, sorting irregular objects, packaging frequently changing products, and supporting workers with physically demanding processes.
The longer-term opportunity is not simply a more capable robotic hand. It is a different deployment model in which manufacturers teach machines through demonstration rather than specifying every movement in advance.
Audi’s involvement gives FLUX-mimic an opportunity to prove that this approach can move beyond controlled research environments. Success on an automotive factory floor would offer an important indication that generative video models can become more than tools for creating digital media, serving instead as part of the intelligence layer that controls machines in the physical world.












