Partnerships
NVIDIA Details Skild AI Collaboration Behind S1 Robot Foundation Model

NVIDIA on September 10, 2026 detailed how Skild AI built its S1 robot foundation model on NVIDIA AI infrastructure, and said the robotics company reached a $100 million annual revenue run rate 10 months after its first commercial deployment.
In the NVIDIA blog post, the company said Skild built S1 and conducted the underlying research on its infrastructure as part of a broader collaboration spanning synthetic data generation, model training, simulation and real-world physical AI deployment. “Learning by experience, and not preprogramming, is the step change that has happened in robotics,” said Deepak Pathak, cofounder and CEO of Skild AI, adding that NVIDIA Isaac Lab and NVIDIA Cosmos help Skild create the scalable, diverse experience its robots need to learn across many scenarios and embodiments. The companies said they are working to move adaptable robot intelligence from the lab into factories and other dynamic operating environments.
Learning Unseen Tasks From One Video
Skild introduced S1 in an August 18, 2026 research post, describing it as a robotic foundation model built from the ground up as an in-context learner. The model takes a video demonstration of a task as input and executes it without updating its weights or undergoing task-specific post-training. Skild describes S1 as the first robotics foundation model to show in-context learning on long-horizon tasks, running up to 10 minutes, that were never seen during pretraining.
An operator records a video of the desired task and provides it to the model as a prompt. The model interprets the demonstrated intent, objects and sequence, then maps them into actions for the robot in front of it, with no retraining, and often for a task not covered by its pretraining dataset. According to both companies, S1 can perform unfamiliar tasks including plant potting, pancake making, pour-over coffee brewing and kit assembly, work that spans dozens of manipulation steps and requires composing skills in sequences the model has not previously performed.
In one plant-potting test logged in Skild’s research post, soil, a pot, a watering can and a plant arrived at the office at 8:54 PM, recording began at 9:16 PM, one egocentric human video demonstration was recorded at 9:22 PM, and S1 began executing the task autonomously on hardware at 9:27 PM. Skild states the time from demonstration to autonomous execution was 11 minutes.
Skild reports that S1 can adjust when objects are moved mid-task, recover from errors, and substitute objects of matched affordance, in one case using a cup of water when the demonstration showed a watering can. The company also reports that the model sometimes improves on flawed demonstrations, treating the demonstration as a specification of the goal rather than a trajectory to reproduce.
Reported Results Against Language-Prompted Policies
NVIDIA’s post reports that in Skild’s tests on new, multistep tasks, S1 succeeded about 66% of the time at each step, compared with 9% for a similar AI system, a gap NVIDIA characterized as a more than sevenfold improvement. Skild’s research post describes the comparison as a controlled study against a language-prompted VLA policy, which receives its task specification in language rather than through demonstration. Both policies were trained on identical data, architectures and compute across pretraining datasets from 1,000 to 100,000 hours, and the 66% and 9% figures were recorded at 100,000 hours on unseen long-horizon tasks, according to Skild.
On tasks seen during pretraining, Skild reports that in-context learning reached 96% success at the largest scale, while at 1,000 hours the language-conditioned policy scored 53% against 43% for in-context learning. Skild also evaluated robustness across five levels of distribution shift, reporting that under the most severe condition, where half of the robot’s actions must be executed with the opposite arm, the language-prompted policy degraded up to three times as much as the in-context policy.
On demonstration efficiency, Skild estimates that a single in-context video matches roughly 380 post-training episodes, a figure it says was interpolated between measured points, and reports that collecting 380 long-horizon demonstrations took 50 to 100 hours of teleoperation. The post-trained baseline eventually reached 86% success with 2,000 demonstrations, according to Skild.
Commercial Traction and the Foxconn Deployment
NVIDIA’s post states that Skild has built more than 60 deployment partnerships, with work spanning manufacturing, logistics, inspection, security, food preparation and other applications.
On the factory floor, Skild, NVIDIA and Foxconn are deploying the Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems. In one demonstrated workflow, a robot installs a busbar and a limit block, fastens 16 screws and adapts to disturbances across the multistep task. NVIDIA’s post states the work requires precise motion, contact-aware control, sequence tracking and recovery when the scene differs from the plan.
Skild first announced the Blackwell production-line plan in a March 19, 2026 post that also disclosed partnerships with ABB Robotics, Universal Robots and Mobile Industrial Robots. In that announcement, Skild described the assembly sequence as a real task performed by humans on a Foxconn line, ending with removal of the limit block, and said its brain was fine-tuned with a small amount of robot data for the workflow.
The NVIDIA Stack Behind S1
According to NVIDIA, its Cosmos open world foundation models help Skild diversify training data and turn video into structured descriptions, while Cosmos Curator helps annotate, filter and organize data at scale. NVIDIA Omniverse libraries and the Isaac Sim framework provide physically based virtual environments for generating data, testing edge cases and validating behaviors before real-world deployment.
Skild uses reinforcement learning in Isaac Lab, an open modular robot learning framework powered by the Newton physics engine, to model physical parameters such as forces, contact, collision and pressure and reduce the simulation-to-reality gap. As models move toward production, NVIDIA Nsight tools help engineers find performance bottlenecks during training, and the TensorRT software development kit optimizes inference so robots can respond quickly in the physical world.
Skild and NVIDIA are also jointly developing GPU-accelerated simulation solvers that model how robots physically touch, grip and manipulate solid objects. The companies said the solvers will soon be made available to all developers as part of Newton.












