Key Moments

Recursive Language Models — Alex Zhang, MIT PhD

Latent Space PodcastLatent Space Podcast
Science & Technology7 min read104 min video
Oct 2, 2026|1,518 views|35|2
Save to Pod
TL;DR

Recursive Language Models (RLMs) offer a new framework for AI agents, but their effectiveness relies on sophisticated harness design and training, not just raw model capability. The true innovation lies in how agents are guided, not just the underlying LLM.

Key Insights

1

A full-scale simulation of OpenAI's approach cost an estimated $40 million, using 10,000 agents over 88 hours to process 130 billion tokens, highlighting the immense computational resources required for complex AI agent systems.

2

Recursive Language Models (RLMs) are presented as a distinct framework where the codebase itself acts as the sole tool, with agents able to call themselves or other functions, storing context in a code environment like a file system.

3

The 'harness tax' paper suggests many harness design choices are inconsequential, implying that the true differentiation may lie in the training data and methodology, rather than the specific harness architecture.

4

While current frontier models can perform complex tasks, they are not yet consistent or robust enough for month-long missions, indicating a 'skill problem' where better system design and prompting could unlock latent capabilities.

5

The concept of 'open-endedness' in AI research, where agents explore problems without predefined goals, is contrasted with goal-directed AI, suggesting the former might be more suitable for fundamental scientific discovery.

6

The effectiveness of AI agents, particularly in complex tasks like solving mathematical problems or navigating codebases, is heavily dependent on their ability to generalize and leverage learned strategies across different lengths and types of problems.

The limitations of current AI agents and the potential of RLMs

The discussion opens by highlighting the perceived similarity between various coding agents like Claude Code, Codex, and By. Alex Zhang posits that while these models possess significant capabilities, they are often wrapped in primitive systems, leading to a 'skill problem.' He argues that even leading models struggle with consistent, long-term task performance over a month. The sheer scale of computation needed for advanced AI systems is illustrated by a hypothetical OpenAI simulation costing $40 million, involving 10,000 agents and 130 billion tokens. This vast expenditure underscores the economic and computational barriers to solving complex problems with current AI. Zhang expresses a desire to see academic research focus on improving these systems, even with existing powerful models like Astra, suggesting that better harness design could unlock latent capabilities.

The evolution of GPU programming communities and benchmarks

Zhang recounts his early involvement with the 'GPU mode' Discord server, which began as a CUDA-focused community and evolved into a broader platform. Initially centered on learning GPU kernel programming and competitive 'code golf,' the community fostered the idea of automating GPU kernel development through large-scale data, akin to platforms like Codeforces. This led to the development of 'Kernel Bench,' aiming to use large language models (LLMs) to automate GPU kernel code generation. He notes the increasing prevalence of AI-generated solutions in benchmarks like the GPU mode leaderboard, but also points out that human expertise, like that of user 'Gurst,' still provides more robust and stable kernels. The conversation touches on the challenge of verifying GPU kernels and the 'reward hacking' often seen in AI-generated code, contrasting it with the need for genuine problem-solving skills.

Rethinking harness design for enhanced AI agent capabilities

A central theme is the critique of current AI agent harnesses, which Zhang believes are too similar and lack innovation. He argues that the focus on 'next token prediction' is an awkward paradigm for many complex tasks, suggesting that harnesses should be designed to augment LLMs' capabilities more effectively. Recursive Language Models (RLMs) are presented as a potential solution, where the codebase itself becomes the tool, and agents can call themselves or other functions within a code environment. This framework stores context in a file system, allowing agents to retain and access their history, which is crucial for long-term tasks. Zhang highlights that RLMs, by structuring communication through code, leverage LLMs' proficiency in code generation for more systematic problem-solving. The RLM's ability to generalize across different tasks and lengths, by recognizing underlying solution strategies, is a key advantage. This generalization capability means training on shorter tasks can yield proficiency in longer ones, a property that traditional methods struggle with.

The 'harness tax' and the debate over harness design importance

The discussion delves into the concept of a 'harness tax,' referring to the overhead and design choices associated with AI agent frameworks. Zhang discusses a paper suggesting that many harness decisions are inconsequential, implying that the focus should shift from specific architectures to the underlying training data and methodology. He notes that companies like OpenAI and Anthropic likely train their models exclusively on their proprietary harnesses. While open-source models can adapt to different harnesses, the frontier models' performance might be tied to their specific training environments. The conversation also touches upon the cost and complexity of these harnesses, questioning whether novel harness designs are truly necessary or if existing frameworks can be improved through better training and optimization. The 'harness tax' paper suggests that beyond user experience, the configuration of agents for effective problem-solving is paramount.

Jeb, RLMs, and the exploration of novel LLM architectures

The introduction of 'Jeb,' a new model architecture, sparks a debate about its significance. Zhang clarifies that while some dismiss it as rehashed concepts, its true value lies in questioning the standard 'text-to-text' design space of LLMs. Jeb explores different output spaces, offering trade-offs between inference speed and accuracy, particularly for tasks with prior user assumptions. This challenges the notion of LLMs solely as auto-regressive decoders. Similarly, RLMs are defended against claims of not being 'language models' by emphasizing that a language model is fundamentally about language modeling, not necessarily a transformer decoder. The discussion highlights how these new architectures open up new research questions about output spaces, inference response times, and the potential for integrating symbolic reasoning within LLMs.

The role of academic research and the 'big bets' in AI

Zhang emphasizes the unique position of PhD students to pursue 'big bets' in research, even on ideas that might seem trivial or uninteresting to industry. He contrasts this with industry labs that focus on immediate applications and benchmarks. The importance of academic research lies in exploring fundamental questions and novel architectures, such as RLMs or 'Swe-bench,' which might not have immediate commercial appeal but can lead to significant breakthroughs later. He cites examples like 'STaR' and 'Quiet-STaR' as simple yet profound ideas that shaped the field. The discussion touches on the 'incentive problem' in academia, where researchers might be pushed towards industry-aligned projects, and advocates for pursuing unconventional ideas that could redefine the field, even if many such bets fail.

Open-endedness versus goal-directed AI and the future of research

The concept of 'open-endedness' in AI research is explored, contrasting it with goal-directed approaches. Open-ended research, akin to exploring problems without a predefined solution, is seen as crucial for fundamental scientific discovery. Zhang compares this to evolutionary research and agent swarms, where the objective is to discover interesting phenomena rather than solve a specific problem. He questions the value of open-ended problems without clear goals, suggesting that alignment with scientific frontiers or targeted problem-solving might be more fruitful. However, he acknowledges that the absence of a defined goal can be a unique form of exploration, potentially uncovering novel approaches or 'hidden gems' within vast amounts of data. This contrasts with the typical machine learning paradigm of optimizing for a specific objective function.

The potential of agent swarms and the challenges of scaling

Agent swarms are discussed as a promising direction, particularly for complex, multi-step problems. The $40 million OpenAI simulation serves as an example of the scale and cost involved. However, challenges remain in designing effective swarm architectures, ensuring efficient communication, and managing computational resources. The 'Cursor' example illustrates a hierarchical structure for agent swarms, which can lead to bottlenecks. Zhang suggests that the ideal swarm would be more decentralized and independent. He also points out that while swarms offer the potential to explore vast problem spaces, discerning useful insights from the noise (token consumption) remains a significant challenge, with many agents potentially being redundant or inefficient. The conversation highlights that the effectiveness of swarms is not guaranteed and depends heavily on intelligent design and targeted training.

Common Questions

GPU Mode is a Discord server community that originated as CUDA Mode, focused on teaching and practicing GPU kernel programming. It hosts lectures and competitions to scale GPU kernel development and automation, which is crucial for researchers.

Topics

Mentioned in this video

Software & Apps
Terra

Mentioned as a potential LLM whose costs could be reduced by more efficient kernels.

Quinn

An open-source model good at using open code, highlighting the need for training within harnesses.

Python

A programming language whose REPL environment can be used to run tools within RLMs.

Alpha Evolved

A system that uses evolutionary search for open-ended problems.

PreFable

An earlier model before Fable, mentioned in the context of RLM performance.

Popcorn

An idea from Marc Sarrazin that became the leaderboard for GPU Mode competitions, focusing on GPU kernel optimization.

Kernel Bench

A benchmark or competition platform that emerged from GPU Mode, exploring the automation of GPU kernel code generation using LLMs.

Gemini

Google's AI model, specifically mentioned in the context of playing Pokemon.

Headlong

A system for 'continuous thinking' by Andy Quynsk that uses an RLM abstraction, formerly called Auto Terminus.

Cursor

An AI code editor, mentioned in relation to multi-agent swarm architecture.

Ocaml

A functional programming language suitable for static analysis and predictive code.

GPT-5.6

A hypothetical future version of GPT models, mentioned in the context of writing more efficient kernels.

Thinky

An AI model or framework praised for breaking the traditional sequence-to-sequence decoder-only paradigm.

Devin

An AI code generation tool, whose emergence made Swebench relevant.

Tufa

An example of a system influenced by RLM abstraction in Kaggle competitions.

Fugu

A system mentioned in the context of open-ended systems, related to model steering.

Haskell

A functional programming language better suited for static analysis and predictive code.

Effect-TS

A functional programming library for TypeScript, offering capabilities for predictive code.

Rexus

A software or system Alex Zhang found boring during an internship.

PyTorch

An open-source machine learning framework; Alex Zhang mentioned it as a past employment path.

Luna

Mentioned as a potential LLM whose costs could be reduced by more efficient kernels.

Prime Agent

A specific RLM (Recursive Language Model) implementation based on Pi, designed with IPython as its sole tool and continuous harness capabilities.

Spark

An example of a generic, unopinionated harness used for coding.

Bash

A Unix shell, whose REPL environment can be used to run scripts within RLMs.

AlphaGeometry

A system that excelled in competitive mathematics, mentioned in relation to Gemini's early work.

GPT-6

A hypothetical next-generation GPT model, used to illustrate that academic papers don't aim for such direct product releases.

Grockbot

An example of a harness that is different in design from others, according to Alex Zhang.

Lisp

A functional programming language suitable for static analysis and predictive code.

Typescript

A programming language, for which Effect-TS provides functional programming capabilities.

GPT

A family of large language models, mentioned as frontier models.

Pi Mono

A simplified version of the Pi framework, used as a basic harness.

LeetCode

A platform for coding challenges, used as an analogy for 'Leak GPU'.

IPython

An interactive Python shell, used as the sole tool within the Prime Agent harness design.

JavaScript

A programming language that makes static analysis for predictive code difficult.

Mamba

An AI model or architecture that released its own kernel, highlighting the need for custom kernel development.

Jev

A language model or system discussed for its novel approach to output space and inference speed, potentially using parallel decoding or diffusion techniques.

Cloud Code

An AI coding assistant, mentioned as a similar harness to Pi and Prime Agent.

Fable

A language model mentioned alongside GPT and GP6 as frontier models.

Claude

An AI model, specifically mentioned for 'Claude plays Pokemon'.

Harvey AI

A legal AI company that successfully trained an RLM for legal work involving document analysis and information retrieval.

Jira

A project management tool, used metaphorically to describe how problems are broken down for agent swarms.

Anti-gravity

A Google product that Alex Zhang found no reason to switch to.

Astra

A powerful frontier model mentioned for its ability to solve complex problems and play games like Kirby.

Gemma

Older open models that were poor at using open code, requiring harness training.

More from Latent Space

View all 259 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free