Key Moments

Robot-Use Agents: Why General-Purpose Models May Win in Robotics

Y CombinatorY Combinator
Science & Technology5 min read30 min video
Sep 26, 2026|59,816 views|423|12
Save to Pod
TL;DR

General-purpose AI models are now controlling robots, performing complex physical tasks with minimal robot-specific training, potentially leading to adaptable robots that can do anything a capable human can.

Key Insights

1

General-purpose AI models, initially excelling at coding, are now demonstrating significant capabilities in controlling diverse robots and learning new physical tasks with little to no robot-specific training.

2

Early research like the RT2 paper fine-tuned language models to output robotic actions (e.g., end-effector poses) instead of language, paving the way for VLA (Vision-Language-Action) models.

3

The 'code as policies' paradigm, exemplified by Voyager, allows AI agents to use programming languages and tools to generate complex robot behaviors, treating code as a form of policy.

4

In-context learning (ICL) with Large Language Models (LLMs) shows rapid adaptation but hits a saturation point after about 20-40 examples, limiting its scalability compared to weight-based learning.

5

Data from computer usage, such as manipulating objects in 3D design software like Blender, is proving crucial for LLMs to develop spatial intelligence and understand physical manipulation for robotics.

6

There's a growing consensus that general-purpose robots, capable of executing any natural language instruction akin to a skilled human, could emerge within the next two years, a development the wider community may not be fully prepared for.

The unexpected generalization of AI to physical tasks

A significant surprise in AI has been the generalization of coding agents beyond software, extending now to robotics. This shift has led researchers and thinkers like MIT professor Philip Isola to propose that we are entering an era of 'robot-use agents.' These are general-purpose AI models capable of controlling various robots, writing new policies, and learning physical tasks with minimal or no robot-specific training. Startups like Waddle Labs and RoboCurve are at the forefront of this movement, using large language models (LLMs) to enhance robot capabilities. Waddle Labs focuses on building LLM-controlled robotic systems by developing effective systems and collecting data for better model training. RoboCurve, on the other hand, specializes in physical AI, evaluating various models and robotic platforms, including LLMs, across different environments and robot types like hands, grippers, and humanoid robots.

From language models to robot control: Early breakthroughs

Early research laid the groundwork for using AI in robotics, notably with the RT2 paper. This work adapted a language model, pre-trained on web text and images, to control robots. Instead of generating English text, the model was fine-tuned to output 'end-effector poses' – specific coordinates that translate into robot joint commands. This approach is similar to how LLMs process information, but RT2 required fine-tuning. The concept evolved towards models being capable enough to perform these tasks directly out-of-the-box. This mirrors the development in LLMs where 'chain-of-thought' prompting allows models to show their reasoning before providing an answer, a capability that wasn't initially present and required explicit enabling. For robots, this means moving beyond direct action outputs to allowing models to 'think' or plan before executing a physical action.

The 'code as policies' paradigm and its implications

The 'code as policies' paradigm represents a significant shift in how AI agents can control robots. Projects like Voyager demonstrated how programming agents could use tools, including Python functions, to generate new capabilities on the fly. For instance, an agent could use basic Python functions (like 'if,' 'for,' 'while') to construct a new tool, such as a custom Python script, which it could then use for better gameplay in Minecraft. This approach treats code generation as a form of policy-making. The insight here is that if AI can master programming, it can potentially automate complex tasks, including writing policies for robots or generating art through JavaScript. This suggests a future where general-purpose programming agents could replace specialized robot controllers or even creative roles by writing the necessary code.

Learning frameworks: In-context learning versus weight updates

Understanding how AI learns is crucial for developing advanced robot-use agents. Two primary modes of learning are in-context learning (ICL) and learning via weight updates (e.g., fine-tuning or full training). ICL involves providing examples within the prompt to guide the model's response, similar to how humans learn from examples. This method is rapid and computationally cheap, requiring no gradient descent. However, ICL has limitations: its effectiveness saturates after a relatively small number of examples (around 20-40), and it's constrained by the model's context window length. Weight-based learning, on the other hand, involves updating the model's internal parameters. While more computationally intensive, it allows for deeper learning and potentially greater performance ceilings, especially with large datasets like those used by companies like Tesla for autonomous driving or Figure for robotics. The challenge lies in efficiently integrating these learning methods.

The role of 'Platonic Representations' and spatial intelligence

The 'Platonic Representation Hypothesis,' inspired by Plato's allegory of the cave, suggests that as AI models are trained on more diverse data, they converge towards consistent, abstract representations of the world. If this hypothesis holds true, powerful LLMs would inherently possess strong representations of the physical world, making them inherently capable of robot control without explicit robot-specific training. This explains why models like OpenAI's Astra, despite being general-purpose LLMs, show remarkable spatial intelligence. Their pre-training on vast amounts of computer usage data, including manipulating objects in 3D software like Blender, helps them develop an understanding of spatial concepts such as 'up,' 'down,' 'left,' and 'right,' which are vital for physical manipulation. This data, while not directly from robotics, provides a foundational understanding of how to interact with and reason about 3D environments, bridging the gap to physical tasks.

Future of robotics: General-purpose robots on the horizon

The convergence of advanced AI models and diverse training data is accelerating progress towards general-purpose robots. There is a growing consensus that within the next two years, we might see robots capable of performing any task given in natural language, much like a competent human. This capability will allow robots to generalize to unseen tasks and environments, marking a 'ChatGPT moment' for robotics. Companies like Waddle Labs are focusing on challenges such as latency and efficient skill integration to bridge the gap between current AI capabilities and the demands of real-world robotic applications. The path forward involves integrating LLM-driven planning with faster, more efficient learned policies, and developing methods for compressing daily experiences into trainable weights, mirroring biological processes like sleep for memory consolidation.

Leveraging General-Purpose Models for Robotics

Practical takeaways from this episode

Do This

Utilize pre-trained language models (LLMs) for robot control, leveraging their broad understanding of data.
Focus on the quality and type of data used for training, as it's crucial for model performance (e.g., computer usage data for spatial reasoning).
Explore 'code as policies' approaches, where LLMs generate code to control robots, enabling them to use tools and create new functions.
Integrate diverse data sources like programming code, computer usage, and first-person perspective videos to build more capable robotic agents.
Consider meta-learning strategies where larger models program smaller, specialized models for specific tasks.
Embrace the 'Platonic Representation Hypothesis' and train models on vast datasets to converge towards consistent world representations.
Optimize for faster response times by compressing learned skills into faster, reusable policies or tools.

Avoid This

Avoid relying solely on robot-specific data if general-purpose LLMs can offer better generalization.
Do not neglect the importance of in-context learning, but be aware of its limitations and saturation points.
Avoid expecting LLMs to perform complex spatial tasks without appropriate training data (e.g., CAD data, computer usage data).
Don't overlook the potential of making robot control interfaces resemble familiar computer usage paradigms.
Avoid solely focusing on fine-tuning for direct action output; explore ways for models to 'think' and generate code/policies.

Common Questions

A robot-use agent is an AI system, often powered by large language models (LLMs), that can understand and execute natural language commands to control robots. General-purpose models are becoming increasingly capable of controlling diverse robots because they leverage broad training data, enabling them to generalize to new tasks and environments without task-specific fine-tuning.

Topics

Mentioned in this video

More from Y Combinator

View all 633 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free