Key Moments
Robot-Use Agents: Why General-Purpose Models May Win in Robotics
Key Moments
General-purpose AI models are now controlling robots, performing complex physical tasks with minimal robot-specific training, potentially leading to adaptable robots that can do anything a capable human can.
Key Insights
General-purpose AI models, initially excelling at coding, are now demonstrating significant capabilities in controlling diverse robots and learning new physical tasks with little to no robot-specific training.
Early research like the RT2 paper fine-tuned language models to output robotic actions (e.g., end-effector poses) instead of language, paving the way for VLA (Vision-Language-Action) models.
The 'code as policies' paradigm, exemplified by Voyager, allows AI agents to use programming languages and tools to generate complex robot behaviors, treating code as a form of policy.
In-context learning (ICL) with Large Language Models (LLMs) shows rapid adaptation but hits a saturation point after about 20-40 examples, limiting its scalability compared to weight-based learning.
Data from computer usage, such as manipulating objects in 3D design software like Blender, is proving crucial for LLMs to develop spatial intelligence and understand physical manipulation for robotics.
There's a growing consensus that general-purpose robots, capable of executing any natural language instruction akin to a skilled human, could emerge within the next two years, a development the wider community may not be fully prepared for.
The unexpected generalization of AI to physical tasks
A significant surprise in AI has been the generalization of coding agents beyond software, extending now to robotics. This shift has led researchers and thinkers like MIT professor Philip Isola to propose that we are entering an era of 'robot-use agents.' These are general-purpose AI models capable of controlling various robots, writing new policies, and learning physical tasks with minimal or no robot-specific training. Startups like Waddle Labs and RoboCurve are at the forefront of this movement, using large language models (LLMs) to enhance robot capabilities. Waddle Labs focuses on building LLM-controlled robotic systems by developing effective systems and collecting data for better model training. RoboCurve, on the other hand, specializes in physical AI, evaluating various models and robotic platforms, including LLMs, across different environments and robot types like hands, grippers, and humanoid robots.
From language models to robot control: Early breakthroughs
Early research laid the groundwork for using AI in robotics, notably with the RT2 paper. This work adapted a language model, pre-trained on web text and images, to control robots. Instead of generating English text, the model was fine-tuned to output 'end-effector poses' – specific coordinates that translate into robot joint commands. This approach is similar to how LLMs process information, but RT2 required fine-tuning. The concept evolved towards models being capable enough to perform these tasks directly out-of-the-box. This mirrors the development in LLMs where 'chain-of-thought' prompting allows models to show their reasoning before providing an answer, a capability that wasn't initially present and required explicit enabling. For robots, this means moving beyond direct action outputs to allowing models to 'think' or plan before executing a physical action.
The 'code as policies' paradigm and its implications
The 'code as policies' paradigm represents a significant shift in how AI agents can control robots. Projects like Voyager demonstrated how programming agents could use tools, including Python functions, to generate new capabilities on the fly. For instance, an agent could use basic Python functions (like 'if,' 'for,' 'while') to construct a new tool, such as a custom Python script, which it could then use for better gameplay in Minecraft. This approach treats code generation as a form of policy-making. The insight here is that if AI can master programming, it can potentially automate complex tasks, including writing policies for robots or generating art through JavaScript. This suggests a future where general-purpose programming agents could replace specialized robot controllers or even creative roles by writing the necessary code.
Learning frameworks: In-context learning versus weight updates
Understanding how AI learns is crucial for developing advanced robot-use agents. Two primary modes of learning are in-context learning (ICL) and learning via weight updates (e.g., fine-tuning or full training). ICL involves providing examples within the prompt to guide the model's response, similar to how humans learn from examples. This method is rapid and computationally cheap, requiring no gradient descent. However, ICL has limitations: its effectiveness saturates after a relatively small number of examples (around 20-40), and it's constrained by the model's context window length. Weight-based learning, on the other hand, involves updating the model's internal parameters. While more computationally intensive, it allows for deeper learning and potentially greater performance ceilings, especially with large datasets like those used by companies like Tesla for autonomous driving or Figure for robotics. The challenge lies in efficiently integrating these learning methods.
The role of 'Platonic Representations' and spatial intelligence
The 'Platonic Representation Hypothesis,' inspired by Plato's allegory of the cave, suggests that as AI models are trained on more diverse data, they converge towards consistent, abstract representations of the world. If this hypothesis holds true, powerful LLMs would inherently possess strong representations of the physical world, making them inherently capable of robot control without explicit robot-specific training. This explains why models like OpenAI's Astra, despite being general-purpose LLMs, show remarkable spatial intelligence. Their pre-training on vast amounts of computer usage data, including manipulating objects in 3D software like Blender, helps them develop an understanding of spatial concepts such as 'up,' 'down,' 'left,' and 'right,' which are vital for physical manipulation. This data, while not directly from robotics, provides a foundational understanding of how to interact with and reason about 3D environments, bridging the gap to physical tasks.
Future of robotics: General-purpose robots on the horizon
The convergence of advanced AI models and diverse training data is accelerating progress towards general-purpose robots. There is a growing consensus that within the next two years, we might see robots capable of performing any task given in natural language, much like a competent human. This capability will allow robots to generalize to unseen tasks and environments, marking a 'ChatGPT moment' for robotics. Companies like Waddle Labs are focusing on challenges such as latency and efficient skill integration to bridge the gap between current AI capabilities and the demands of real-world robotic applications. The path forward involves integrating LLM-driven planning with faster, more efficient learned policies, and developing methods for compressing daily experiences into trainable weights, mirroring biological processes like sleep for memory consolidation.
Mentioned in This Episode
●Software & Apps
●Companies
●Organizations
●Studies Cited
●Concepts
●People Referenced
Leveraging General-Purpose Models for Robotics
Practical takeaways from this episode
Do This
Avoid This
Common Questions
A robot-use agent is an AI system, often powered by large language models (LLMs), that can understand and execute natural language commands to control robots. General-purpose models are becoming increasingly capable of controlling diverse robots because they leverage broad training data, enabling them to generalize to new tasks and environments without task-specific fine-tuning.
Topics
Mentioned in this video
A startup focused on building large language models that control robots by developing effective systems and collecting data for better training.
A company specializing in physical AI, measuring and evaluating various robots, models (including LLMs), and traditional approaches across different environments and robot types.
A company mentioned in the context of impressive few-shot learning performance using in-context learning.
A company whose software was mentioned in the context of Xerox PARC's work, where its graphical interfaces aided in simulating physical world interactions for robot learning.
Mentioned as an example of a company with vast data resources that would not rely solely on in-context learning for critical tasks like autonomous driving.
A startup accelerator that is accepting applications for its next batch.
An allegory used to explain the 'Platonic Representation Hypothesis', referring to seeing shadows of a diverse reality.
A learning method where a model is provided with examples within its prompt to guide its output, commonly used by LLM users.
A measure of the computational resources needed to describe a string, related to the idea of compressed code and efficient representation.
A parameter-efficient fine-tuning technique for large language models, mentioned in the context of different levels of model training.
Referenced as a paradigm shift moment in AI, similar to what is expected for general-purpose robots in the near future.
3D computer graphics software used in demonstrations of Astra's spatial intelligence, where it controls the software to create 3D images.
A framework for data aggregation in reinforcement learning, used as an analogy for how AI systems might compress and train weights during a 'sleep' phase.
A language model mentioned as being used directly for robot control and in demonstrations, capable of writing code to control robot joints and thinking through tasks.
The programming language used for implementing functions and policies in robotic control, as discussed in the context of 'code as policies' research.
A model used to control robotic arms, demonstrated picking up a cube and placing it in a bowl using camera inputs and generating code for robot joint control.
Mentioned as an API or framework that can be used for robot control, offering advantages over direct LLM control for repetitive or previously performed tasks.
A CAD software mentioned in the context of Xerox PARC's work, where its graphical interfaces aided in simulating physical world interactions for robot learning.
A tool mentioned for its library of skills and its ability to restructure them into a more compact representation during a 'sleep' phase, similar to potential robot skill management.
A research center from the 1980s mentioned for its work in making GUIs more like the physical world, inadvertently creating environments beneficial for robot learning.
A research team whose papers on 'code as policies' demonstrated impressive capabilities in creating robotic control functions from Python code.
Mentioned as a source of research related to cloud robotics and improving LLM capabilities in physical tasks by making robot interaction tools resemble computer use.
An MIT professor who suggested we might be entering an era of robot-use agents, where general-purpose models can make robots more capable.
Author of the research paper 'The Platonic Representation Hypothesis', which suggests that language models converge towards consistent world representations with more training data.
More from Y Combinator
View all 633 summaries
37 minThe State of Startups in 2026
61 minWhy The Harness Matters More Than The Model | YC Paper Club
58 minOpen Models Are Collapsing The Cost Of AI
22 minPaul Graham On Startups, Ambition, and Great Founders
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free