Key Moments

Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | The GPU Economy

Stanford OnlineStanford Online
Education5 min read57 min video
Jul 23, 2026|1,836 views|131|5
Save to Pod

Want to know something specific about what's covered?

We've already dissected every moment. Ask and we will deliver (with timestamps).

TL;DR

AI inference costs are plummeting, making AI companies profitable and driving massive demand. While this revolutionizes capabilities, it also presents significant challenges in distribution and societal adaptation.

Key Insights

1

The cost of inference has dropped by 90% in the last year and closer to 99% over two and a half years.

2

NVIDIA has committed to $1 trillion in sales over the next eight quarters, indicating demand far outstrips supply.

3

Anthropic added $10 billion in annualized revenue in March alone, demonstrating a massive shift in customer willingness to pay for AI capabilities.

4

While training and inference costs are decreasing, models are simultaneously becoming exponentially larger, with parameters moving from billions to potential trillions.

5

The fundamental flop calculation for generating a single token is roughly the parameter size of the model times the context length squared.

6

AI is being harnessed to design the next generation of chips, creating a recursive loop of innovation and necessitating continued infrastructure buildout.

The non-zero cost of AI distribution

Unlike traditional software, which has near-zero incremental distribution costs, AI applications require significant compute. This fundamental difference drives the economics of the current AI supercycle. The discussion highlights that while software 'ate the world' due to its easily scalable nature, AI's dependency on compute introduces new economic considerations and constraints that are shaping the industry.

Compute as the atomic unit of AI

The conversation centers on compute as the foundational element for AI, akin to tokens being the atomic unit of intelligence. Sunny Madra from Groq (now NVIDIA) explains that the core of any AI problem, especially token generation, involves an immense amount of mathematical computation. This is why the demand for compute is exploding. The complexity of generating a single token is often described as the parameter size of the model multiplied by the context length squared. This means that as models grow larger and context windows expand, the compute requirements escalate dramatically, placing a premium on efficient hardware and architectures.

The Grock architecture and deterministic inference

Grock, founded by Jonathan Ross (creator of Google's TPU), developed a chip with a data flow architecture that is fully deterministic. This means its compiler predetermines where all calculations will occur. This deterministic nature is crucial for efficient AI inference, as it minimizes wasted computation. Grock's approach, which emphasizes SRAM with high bandwidth over massive compute, proved more efficient for inference tasks than traditional GPUs, especially when integrated with NVIDIA's architecture. This led to an acquisition by NVIDIA for $20 billion, their largest ever, driven by the potential to significantly scale token generation within existing power and footprint constraints.

The inference inferno: exploding token consumption and demand

The transition from pre-training models to inference-time reasoning has led to an astonishing increase in token consumption. Brad Gerstner notes that Jensen Huang of NVIDIA described this shift as potentially a billion-fold increase in demand for inference capabilities. This surge, even before the advent of agents that perform actions, caused existing compute infrastructures to strain. Companies like OpenAI and Anthropic, operating with limited compute budgets (e.g., one gigawatt each), faced critical challenges in meeting this parabolic growth in demand. This underscored the urgent need for more token-efficient models and architectures.

The NVIDIA-Groq partnership: accelerating token output

A pivotal moment discussed is Sunny Madra's idea to partner Groq's deterministic, SRAM-based inference chips with NVIDIA's GPUs via NVLink Fusion. This integration allows parts of the inference process, where Groq's architecture excels, to be offloaded from NVIDIA's compute-heavy GPUs. The result is a significant increase in efficiency: the same power footprint can generate two and a half times more tokens. This collaboration was critical for NVIDIA, as it complemented their existing ecosystem and demonstrated a path to greatly enhancing inference output, a key economic driver for AI companies like OpenAI and Anthropic.

Plummeting inference costs and rising profitability

The economic landscape of AI has transformed due to rapidly decreasing inference costs. Over the past year, inference costs have fallen by approximately 90%, and by nearly 99% over two and a half years. This deflationary trend is a byproduct of extreme co-design across the entire compute 'factory,' including chips, packaging, and lithography. Concurrently, as AI capabilities become more valuable and customer willingness to pay increases, AI companies like OpenAI and Anthropic have shifted from negative to highly positive gross margins. This economic viability is crucial for sustaining the massive compute investments required.

The agent revolution and increased value delivery

Beyond generating answers, AI is moving into an 'action' phase with agents that can perform tasks like building applications, resolving customer service issues, or booking travel. While this requires an order of magnitude more tokens to execute, the value delivered to the end consumer increases by 100x. This dramatic increase in perceived value drives higher willingness to pay. Anthropic's recent substantial revenue growth illustrates how reaching a threshold of intelligent capability compels customers to adopt AI solutions, fundamentally changing the economic equation for AI services.

The future of AI: exponential growth, societal challenges, and 'bionic' adaptation

Speakers contend that we are nearing the end of an exponential curve in AI development, with capabilities accelerating faster than anticipated. While this promises an 'age of abundance,' it also presents unprecedented distribution and societal adaptation challenges. The need to manage wealth accumulation and the equitable distribution of AI's benefits is paramount. Individuals are urged to become 'bionic' by leveraging AI, as raw intellectual capability (IQ) becomes commoditized, while emotional intelligence (EQ), teamwork, and persuasion skills become more critical. The ongoing innovation, fueled by AI designing new chips and models, ensures continued progress, but necessitates active engagement with public policy and social contracts to navigate the profound changes AI will bring.

Common Questions

Unlike traditional software, which has near-zero incremental cost of distribution, AI applications require significant compute power as more users engage, making distribution inherently more costly.

Topics

Mentioned in this video

Companies
Altimeter

An investment firm founded by Brad Gerstner, managing over $15 billion. It has invested in numerous tech companies across various market cycles.

OpenAI

A leading AI research and deployment company. Altimeter is a significant investor, and its models like Whisper and ChatGPT were mentioned.

Anthropic

An AI safety and research company. Altimeter is a significant investor, and its models like Claude and Mythos were discussed.

NVIDIA

A dominant player in AI hardware, particularly GPUs. NVIDIA acquired Grok's platform and is a key focus of the discussion regarding compute and economic scale.

Cerebras

A company that developed fast inference chips, competing in the AI hardware space. Brad Gerstner was on its board.

Extreme Labs

A mobile development shop co-founded by Sunny Madra and acquired by Pivotal.

Ford

Automotive company that acquired Autonomic, a smart mobility platform co-founded by Sunny Madra, where he then led their internal innovation lab, Ford X.

Autonomic

A smart mobility platform co-founded by Sunny Madra and acquired by Ford.

Definitive Intelligence

A company acquired by Grok, where Sunny Madra served as president and helped launch Grok Cloud.

Snowflake

Mentioned as a successful investment by Altimeter, it's a cloud-based data warehousing company.

Confluent

Mentioned as a successful investment by Altimeter, it's a company focused on data streaming.

GitLab

Mentioned as a successful investment by Altimeter, it's a web-based DevOps lifecycle tool.

Intel

Mentioned in the context of competing with NVIDIA's AI hardware, specifically with their Gaudi accelerators.

AMD

Mentioned as a competitor to NVIDIA in the AI hardware space, with their Instinct accelerators.

Apple

Discussed regarding its strategy for AI on devices, with concerns about privacy and the potential for its Gemini model to improve Siri.

Google

Mentioned for its past infrastructure work in video and search, and its development of TPUs. Also mentioned in the context of Gemini models.

Microsoft

Mentioned in the context of Mustafa's comments on compute power relative to GPT-2 capabilities, and as a participant in Project Glasswing.

TSMC

Taiwan Semiconductor Manufacturing Company, a primary player in the global semiconductor supply chain, particularly for advanced chip manufacturing.

More from Stanford Online

View all 91 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free