Drift

Research

The Race to Build a Robot Foundation Model

Skild, Dyna, Generalist, NVIDIA and Google DeepMind are taking different approaches to building foundation models for general-purpose robotics.

Drift TeamSep 24, 2026 · 8 min read

Imagine training a robot once, then giving that same brain completely different tasks.

Instead of building one model for folding towels, another for picking objects, and another for navigating a warehouse, the goal is to build a system that can learn a broad understanding of the physical world and apply it across tasks, environments, and eventually different robot bodies.

That's the idea behind robot foundation models, and it has become one of the biggest races in physical AI.

 

The companies working on this problem are taking very different routes. Some are scaling human video. Others are collecting robot demonstrations, building models that learn from a few examples, or combining foundation models with simulation and embodied reasoning. The architectures differ, but the underlying question is similar: how do you build one model that knows enough about the physical world to be useful almost anywhere?

Why Robotics Needs Foundation Models

Most robotic systems have traditionally been built around a particular task and environment. A robot might be trained to pick objects from a conveyor belt, fold a specific type of material, or perform a fixed sequence of movements. That works well when the conditions are controlled, but the physical world is full of variation.

A different object may require a different grasp. A new room may require a different movement strategy. A different robot body changes the relationship between perception and action.

 

Foundation models approach the problem differently. Instead of learning only the exact behavior needed for one task, they are pretrained on much broader data and then adapted to particular robots or tasks. The hope is that the model develops reusable representations of objects, actions, spatial relationships, and physical interactions.

 

This is similar to the shift that happened in language and vision. Large pretrained models gave developers a starting point that could be adapted to many downstream applications instead of training a separate model from scratch every time.

Robotics is attempting something similar, but with a much harder constraint: the output isn't text or an image. It's physical action.

That means the model has to connect what it sees and understands with what a robot can actually do.

Different Companies, Different Bets

There isn't a single recipe for building this kind of model.

Skild AI is pursuing what it calls omni-bodied intelligence: a unified foundation model capable of controlling different kinds of robots and performing different tasks. Its Skild Brain is designed around the idea that data from different robot embodiments can contribute to a shared model, creating a feedback loop in which more deployments generate more data for future versions.

 

More recently, Skild introduced S1, a robotics foundation model built around in-context learning. Instead of requiring fine-tuning for every new behavior, S1 can use a demonstration as context for learning a task. Skild's September 2026 work has also explored physical self-play, where a pretrained model can improve through reinforcement learning in simulation.

 

Dyna Robotics is attacking another major bottleneck: data. Its latest Dyna-2 world-action model was pretrained on more than one million hours of egocentric human video. Dyna reports that increasing the amount of human video produced predictable improvements on held-out robot data, which it describes as a human-to-robot transfer scaling law.

 

The interesting part of Dyna's approach is that it treats human video as a potentially enormous source of physical experience. Rather than relying entirely on expensive robot teleoperation, the model learns from people cooking, folding, cleaning, assembling, and interacting with objects, then transfers some of that knowledge across the embodiment gap.

 

Generalist AI is pushing on another part of the problem: how quickly a robot can acquire a new skill. Its GEN-1.5 model is designed for one-shot and few-shot physical learning, with Generalist reporting that the model can learn a new task from a single demonstration without gradient updates or fine-tuning. The company describes this as in-context learning for physical skills.

 

Then there is NVIDIA, which is taking a broader ecosystem approach around its GR00T family. GR00T N1 is an open foundation model for humanoid robots, combining vision-language understanding with action generation. NVIDIA also surrounds the model with simulation, synthetic-data generation, and robot-learning infrastructure through Isaac. Its later GR00T N1.6 release expanded the model's real-robot evaluation across multiple embodiments.

 

Google DeepMind is taking yet another route with Gemini Robotics. Its models combine the capabilities of Gemini with vision-language-action control and embodied reasoning. Gemini Robotics 2 is designed to control different robot bodies, while its ER models focus on spatial and physical reasoning, planning, and understanding the environment.

 

So while these companies are often discussed under the same "robot foundation model" label, they're not simply building copies of one another. They're making different bets about what kind of data, architecture, and learning process will produce general physical intelligence.

The Real Race Is About Data and Generalization

Building the model is only part of the problem. The bigger question is where its physical knowledge comes from.

 

Robot data is expensive. Someone has to operate the robot, collect demonstrations, reset the environment, label or process the data, and repeat the process across enough objects, tasks, and environments to make the model generalize.

Human video offers scale, but creates an embodiment gap because humans and robots do not have identical bodies or sensors. Simulation offers control and unlimited variation, but simulated physics and visuals don't perfectly match reality.

Robot demonstrations provide direct action data, but are expensive to collect at scale.

That is why the current approaches are starting to look less like competing single techniques and more like different pieces of the same infrastructure.

 

Dyna is demonstrating what happens when human video is scaled to extraordinary volumes. NVIDIA is combining real demonstrations with synthetic data and simulation. Google DeepMind is connecting general multimodal reasoning with robot action. Skild is combining large-scale pretraining with in-context learning and, increasingly, post-training through physical self-play. Generalist is exploring how much new physical behavior can emerge from a small number of demonstrations.

 

This also explains why simulation remains important even in an era of foundation models. A model that has learned from the real world still needs ways to experience situations that are rare, dangerous, or expensive to reproduce physically. Simulation can provide that additional training ground.

 

We've looked at this from another angle in why robots train in simulation before the real world, and the same idea increasingly applies to foundation models: more useful experience can mean more capable robots.

One Brain for Many Robots?

The long-term vision is bigger than making today's robots slightly better.

Imagine a model that has learned from millions of hours of human activity, robot demonstrations, simulation, and interaction with the physical world. You give it a new task, show it a few examples, or simply describe what you want. The model figures out how the task relates to what it already knows and produces actions for the robot in front of it.

 

That is much closer to the way people use intelligence. We don't need to relearn the concept of "picking something up" every time we encounter a new object. We transfer what we already know to the situation in front of us.

 

The robotics industry is still far from that level of generality. Today's foundation models have important limitations, and many capabilities still require task-specific post-training, careful evaluation, or particular robot embodiments. Even the companies building the most general systems are still experimenting with the right combination of data, model architecture, simulation, and real-world learning.

But the direction is becoming clearer.

 

The next major breakthrough in robotics may not come from giving a robot another degree of freedom or making one controller slightly better. It may come from building a model that can take knowledge learned from one situation and carry it into the next.

The race isn't just to build better robots. It's to build the intelligence that can make many different robots more capable.

FAQ

What is a robot foundation model?

A robot foundation model is a broadly pretrained model designed to support multiple robotic tasks, environments, or embodiments rather than being built for only one specific behavior.

How are robot foundation models different from traditional robot controllers?

Traditional controllers are often designed around specific tasks, robot hardware, and environments. Foundation models aim to learn more general representations and behaviors that can be adapted across different situations.

What data are companies using to train these models?

Approaches vary. Current systems use combinations of human video, robot demonstrations, teleoperation data, simulation, synthetic data, and other multimodal datasets. Dyna, for example, reports training Dyna-2 on more than one million hours of egocentric human video.

Are these models already general-purpose robots?

Not completely. Current foundation models demonstrate increasingly broad capabilities, but generalization across arbitrary tasks, environments, and robot bodies remains an active research problem.

Why is simulation still important?

Simulation provides a scalable way to generate experiences that can be difficult, expensive, or unsafe to collect on physical robots. NVIDIA, for example, combines its robot foundation models with Isaac simulation and synthetic-data tools.

Related Reading

Enjoyed this one? Send it to someone who’d find it useful.