Key Moments
Recursive Language Models — Alex Zhang, MIT PhD
Key Moments
Recursive Language Models (RLMs) offer a new framework for AI agents, but their effectiveness relies on sophisticated harness design and training, not just raw model capability. The true innovation lies in how agents are guided, not just the underlying LLM.
Key Insights
A full-scale simulation of OpenAI's approach cost an estimated $40 million, using 10,000 agents over 88 hours to process 130 billion tokens, highlighting the immense computational resources required for complex AI agent systems.
Recursive Language Models (RLMs) are presented as a distinct framework where the codebase itself acts as the sole tool, with agents able to call themselves or other functions, storing context in a code environment like a file system.
The 'harness tax' paper suggests many harness design choices are inconsequential, implying that the true differentiation may lie in the training data and methodology, rather than the specific harness architecture.
While current frontier models can perform complex tasks, they are not yet consistent or robust enough for month-long missions, indicating a 'skill problem' where better system design and prompting could unlock latent capabilities.
The concept of 'open-endedness' in AI research, where agents explore problems without predefined goals, is contrasted with goal-directed AI, suggesting the former might be more suitable for fundamental scientific discovery.
The effectiveness of AI agents, particularly in complex tasks like solving mathematical problems or navigating codebases, is heavily dependent on their ability to generalize and leverage learned strategies across different lengths and types of problems.
The limitations of current AI agents and the potential of RLMs
The discussion opens by highlighting the perceived similarity between various coding agents like Claude Code, Codex, and By. Alex Zhang posits that while these models possess significant capabilities, they are often wrapped in primitive systems, leading to a 'skill problem.' He argues that even leading models struggle with consistent, long-term task performance over a month. The sheer scale of computation needed for advanced AI systems is illustrated by a hypothetical OpenAI simulation costing $40 million, involving 10,000 agents and 130 billion tokens. This vast expenditure underscores the economic and computational barriers to solving complex problems with current AI. Zhang expresses a desire to see academic research focus on improving these systems, even with existing powerful models like Astra, suggesting that better harness design could unlock latent capabilities.
The evolution of GPU programming communities and benchmarks
Zhang recounts his early involvement with the 'GPU mode' Discord server, which began as a CUDA-focused community and evolved into a broader platform. Initially centered on learning GPU kernel programming and competitive 'code golf,' the community fostered the idea of automating GPU kernel development through large-scale data, akin to platforms like Codeforces. This led to the development of 'Kernel Bench,' aiming to use large language models (LLMs) to automate GPU kernel code generation. He notes the increasing prevalence of AI-generated solutions in benchmarks like the GPU mode leaderboard, but also points out that human expertise, like that of user 'Gurst,' still provides more robust and stable kernels. The conversation touches on the challenge of verifying GPU kernels and the 'reward hacking' often seen in AI-generated code, contrasting it with the need for genuine problem-solving skills.
Rethinking harness design for enhanced AI agent capabilities
A central theme is the critique of current AI agent harnesses, which Zhang believes are too similar and lack innovation. He argues that the focus on 'next token prediction' is an awkward paradigm for many complex tasks, suggesting that harnesses should be designed to augment LLMs' capabilities more effectively. Recursive Language Models (RLMs) are presented as a potential solution, where the codebase itself becomes the tool, and agents can call themselves or other functions within a code environment. This framework stores context in a file system, allowing agents to retain and access their history, which is crucial for long-term tasks. Zhang highlights that RLMs, by structuring communication through code, leverage LLMs' proficiency in code generation for more systematic problem-solving. The RLM's ability to generalize across different tasks and lengths, by recognizing underlying solution strategies, is a key advantage. This generalization capability means training on shorter tasks can yield proficiency in longer ones, a property that traditional methods struggle with.
The 'harness tax' and the debate over harness design importance
The discussion delves into the concept of a 'harness tax,' referring to the overhead and design choices associated with AI agent frameworks. Zhang discusses a paper suggesting that many harness decisions are inconsequential, implying that the focus should shift from specific architectures to the underlying training data and methodology. He notes that companies like OpenAI and Anthropic likely train their models exclusively on their proprietary harnesses. While open-source models can adapt to different harnesses, the frontier models' performance might be tied to their specific training environments. The conversation also touches upon the cost and complexity of these harnesses, questioning whether novel harness designs are truly necessary or if existing frameworks can be improved through better training and optimization. The 'harness tax' paper suggests that beyond user experience, the configuration of agents for effective problem-solving is paramount.
Jeb, RLMs, and the exploration of novel LLM architectures
The introduction of 'Jeb,' a new model architecture, sparks a debate about its significance. Zhang clarifies that while some dismiss it as rehashed concepts, its true value lies in questioning the standard 'text-to-text' design space of LLMs. Jeb explores different output spaces, offering trade-offs between inference speed and accuracy, particularly for tasks with prior user assumptions. This challenges the notion of LLMs solely as auto-regressive decoders. Similarly, RLMs are defended against claims of not being 'language models' by emphasizing that a language model is fundamentally about language modeling, not necessarily a transformer decoder. The discussion highlights how these new architectures open up new research questions about output spaces, inference response times, and the potential for integrating symbolic reasoning within LLMs.
The role of academic research and the 'big bets' in AI
Zhang emphasizes the unique position of PhD students to pursue 'big bets' in research, even on ideas that might seem trivial or uninteresting to industry. He contrasts this with industry labs that focus on immediate applications and benchmarks. The importance of academic research lies in exploring fundamental questions and novel architectures, such as RLMs or 'Swe-bench,' which might not have immediate commercial appeal but can lead to significant breakthroughs later. He cites examples like 'STaR' and 'Quiet-STaR' as simple yet profound ideas that shaped the field. The discussion touches on the 'incentive problem' in academia, where researchers might be pushed towards industry-aligned projects, and advocates for pursuing unconventional ideas that could redefine the field, even if many such bets fail.
Open-endedness versus goal-directed AI and the future of research
The concept of 'open-endedness' in AI research is explored, contrasting it with goal-directed approaches. Open-ended research, akin to exploring problems without a predefined solution, is seen as crucial for fundamental scientific discovery. Zhang compares this to evolutionary research and agent swarms, where the objective is to discover interesting phenomena rather than solve a specific problem. He questions the value of open-ended problems without clear goals, suggesting that alignment with scientific frontiers or targeted problem-solving might be more fruitful. However, he acknowledges that the absence of a defined goal can be a unique form of exploration, potentially uncovering novel approaches or 'hidden gems' within vast amounts of data. This contrasts with the typical machine learning paradigm of optimizing for a specific objective function.
The potential of agent swarms and the challenges of scaling
Agent swarms are discussed as a promising direction, particularly for complex, multi-step problems. The $40 million OpenAI simulation serves as an example of the scale and cost involved. However, challenges remain in designing effective swarm architectures, ensuring efficient communication, and managing computational resources. The 'Cursor' example illustrates a hierarchical structure for agent swarms, which can lead to bottlenecks. Zhang suggests that the ideal swarm would be more decentralized and independent. He also points out that while swarms offer the potential to explore vast problem spaces, discerning useful insights from the noise (token consumption) remains a significant challenge, with many agents potentially being redundant or inefficient. The conversation highlights that the effectiveness of swarms is not guaranteed and depends heavily on intelligent design and targeted training.
Mentioned in This Episode
●Products
●Software & Apps
●Companies
●Organizations
●Books
●Studies Cited
●Concepts
●People Referenced
Common Questions
GPU Mode is a Discord server community that originated as CUDA Mode, focused on teaching and practicing GPU kernel programming. It hosts lectures and competitions to scale GPU kernel development and automation, which is crucial for researchers.
Topics
Mentioned in this video
MIT PhD student known for his work on Recursive Language Models (RLMs) and involvement with GPU Mode.
The creator of Headlong (formerly Auto Terminus), a system for continuous thinking.
Guest on a previous podcast episode who discussed competitive vs. cooperative agents.
One of the founders of GPU Mode, known for his work and lectures within the community.
Renowned mathematician, mentioned for his skepticism about using AI in mathematics.
Mentioned as one of the founders of GPU Mode alongside Jeremy Howard.
Google AI leader who tracks microjoules in energy consumption for efficiency.
Mentioned as one of the founders of GPU Mode alongside Marc Andreessen.
Researcher behind STaR and Quiet-STaR, whose work Alex Zhang admires for its simplicity and vision.
Co-founder of Sakana AI, described as intelligent with a good vision.
From Hugging Face, discussed model calibration issues.
Co-founder of Recursive Super Intelligence and head of open-ended systems at Google.
Origin of the 'Infinite Attention' research paper.
A Japanese AI company that spun off from GDM, known for its focus on open-ended evolutionary research and culturally specialized models.
A Discord server community focused on learning how to write GPU kernels, which evolved from CUDA Mode and now hosts competitions and lectures. Alex Zhang is involved in it.
A technology company mentioned for its AI contributions.
A platform for code hosting and version control, where RLM resources can be found.
An AI community and platform, where Clementine Foriou managed evaluation.
An example of a generic, unopinionated harness used for coding.
A leading AI research lab, mentioned as a competitor to OpenAI, possessing vast computational resources.
A company collaborating with Harvey AI, involved in continuous thinking systems.
A social media company where Alex Zhang interned and worked on GPU kernel writing related to Infinite Attention.
A company mentioned for its work on open-ended systems.
A leading AI research lab, often contrasted with smaller labs due to its computational resources and strategy.
Mentioned as a potential LLM whose costs could be reduced by more efficient kernels.
An open-source model good at using open code, highlighting the need for training within harnesses.
A programming language whose REPL environment can be used to run tools within RLMs.
A system that uses evolutionary search for open-ended problems.
An earlier model before Fable, mentioned in the context of RLM performance.
An idea from Marc Sarrazin that became the leaderboard for GPU Mode competitions, focusing on GPU kernel optimization.
A benchmark or competition platform that emerged from GPU Mode, exploring the automation of GPU kernel code generation using LLMs.
Google's AI model, specifically mentioned in the context of playing Pokemon.
A system for 'continuous thinking' by Andy Quynsk that uses an RLM abstraction, formerly called Auto Terminus.
An AI code editor, mentioned in relation to multi-agent swarm architecture.
A functional programming language suitable for static analysis and predictive code.
A hypothetical future version of GPT models, mentioned in the context of writing more efficient kernels.
An AI model or framework praised for breaking the traditional sequence-to-sequence decoder-only paradigm.
An AI code generation tool, whose emergence made Swebench relevant.
An example of a system influenced by RLM abstraction in Kaggle competitions.
A system mentioned in the context of open-ended systems, related to model steering.
A functional programming language better suited for static analysis and predictive code.
A functional programming library for TypeScript, offering capabilities for predictive code.
A software or system Alex Zhang found boring during an internship.
An open-source machine learning framework; Alex Zhang mentioned it as a past employment path.
Mentioned as a potential LLM whose costs could be reduced by more efficient kernels.
A specific RLM (Recursive Language Model) implementation based on Pi, designed with IPython as its sole tool and continuous harness capabilities.
An example of a generic, unopinionated harness used for coding.
A Unix shell, whose REPL environment can be used to run scripts within RLMs.
A system that excelled in competitive mathematics, mentioned in relation to Gemini's early work.
A hypothetical next-generation GPT model, used to illustrate that academic papers don't aim for such direct product releases.
An example of a harness that is different in design from others, according to Alex Zhang.
A functional programming language suitable for static analysis and predictive code.
A programming language, for which Effect-TS provides functional programming capabilities.
A family of large language models, mentioned as frontier models.
A simplified version of the Pi framework, used as a basic harness.
A platform for coding challenges, used as an analogy for 'Leak GPU'.
An interactive Python shell, used as the sole tool within the Prime Agent harness design.
A programming language that makes static analysis for predictive code difficult.
An AI model or architecture that released its own kernel, highlighting the need for custom kernel development.
A language model or system discussed for its novel approach to output space and inference speed, potentially using parallel decoding or diffusion techniques.
An AI coding assistant, mentioned as a similar harness to Pi and Prime Agent.
A language model mentioned alongside GPT and GP6 as frontier models.
An AI model, specifically mentioned for 'Claude plays Pokemon'.
A legal AI company that successfully trained an RLM for legal work involving document analysis and information retrieval.
A project management tool, used metaphorically to describe how problems are broken down for agent swarms.
A Google product that Alex Zhang found no reason to switch to.
A powerful frontier model mentioned for its ability to solve complex problems and play games like Kirby.
Older open models that were poor at using open code, requiring harness training.
A prompting technique for LLMs, mentioned as a simple yet impactful idea.
An extension of STaR, also by Eric Zelikman, contributing to foundational ideas in AI.
A prompting framework combining reasoning and acting, mentioned for its perceived simplicity.
The original name for GPU Mode, focused specifically on NVIDIA's CUDA platform.
A concept from Eric Zelikman's work, which Alex Zhang considers simple yet foundational for the field's direction.
Complex mathematical equations that OpenAI claimed to solve by directing a model towards them.
A benchmark suite for machine learning performance, compared to GPU Mode's community-driven approach.
A research concept developed by Chemy, mentioned for its ingenious nature.
A technique for optimizing attention mechanisms in transformers, which impressed Alex Zhang.
The linguistic hypothesis that the structure of a language affects its speakers' world view or cognition, applied to how LLMs think in different languages.
A video game franchise, specifically 'Claude plays Pokemon' was an old benchmark for AI.
A film cited as an example of aliens ('heptapods') experiencing time non-linearly, used to illustrate a contrast with auto-regressive language models.
A video game mentioned as an example of a game that the Astra model can solve.
A video game example used to illustrate LLM capabilities in gaming.
A competitive programming website used as an analogy for scaling GPU kernel development.
A platform for data science and machine learning competitions, hosting the ARC AGI 3 competition.
Alex Zhang's affiliation as a PhD student and where his advisor Omer works.
A leading AI research lab, mentioned for its substantial resources and sometimes criticized for bureaucracy.
A research paper suggesting that many harness designs are largely similar and that differences in LLM performance within them are primarily cost-related.
A research paper from Google that Alex Zhang attempted to implement with specialized kernels during his Snapchat internship.
More from Latent Space
View all 259 summaries
41 minInside OpenAI DevDay: Superhuman Computer Use, Decisions API, and the AI Cloud — Ari & Nikunj
95 minThe Future of Claude Code: Mods, Mutable Software, & Multiplayer Agents — Thariq Shihipar, Anthropic
99 minRunway’s Bet Beyond Video: World Models, Robotics, and the Neural OS — Anastasis Germanidis
83 minThe $10 Trillion Token Economy — Alex Atallah, OpenRouter & Anjney Midha, AMP
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free