Key Moments
Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | The GPU Economy
Want to know something specific about what's covered?
We've already dissected every moment. Ask and we will deliver (with timestamps).
Key Moments
AI inference costs are plummeting, making AI companies profitable and driving massive demand. While this revolutionizes capabilities, it also presents significant challenges in distribution and societal adaptation.
Key Insights
The cost of inference has dropped by 90% in the last year and closer to 99% over two and a half years.
NVIDIA has committed to $1 trillion in sales over the next eight quarters, indicating demand far outstrips supply.
Anthropic added $10 billion in annualized revenue in March alone, demonstrating a massive shift in customer willingness to pay for AI capabilities.
While training and inference costs are decreasing, models are simultaneously becoming exponentially larger, with parameters moving from billions to potential trillions.
The fundamental flop calculation for generating a single token is roughly the parameter size of the model times the context length squared.
AI is being harnessed to design the next generation of chips, creating a recursive loop of innovation and necessitating continued infrastructure buildout.
The non-zero cost of AI distribution
Unlike traditional software, which has near-zero incremental distribution costs, AI applications require significant compute. This fundamental difference drives the economics of the current AI supercycle. The discussion highlights that while software 'ate the world' due to its easily scalable nature, AI's dependency on compute introduces new economic considerations and constraints that are shaping the industry.
Compute as the atomic unit of AI
The conversation centers on compute as the foundational element for AI, akin to tokens being the atomic unit of intelligence. Sunny Madra from Groq (now NVIDIA) explains that the core of any AI problem, especially token generation, involves an immense amount of mathematical computation. This is why the demand for compute is exploding. The complexity of generating a single token is often described as the parameter size of the model multiplied by the context length squared. This means that as models grow larger and context windows expand, the compute requirements escalate dramatically, placing a premium on efficient hardware and architectures.
The Grock architecture and deterministic inference
Grock, founded by Jonathan Ross (creator of Google's TPU), developed a chip with a data flow architecture that is fully deterministic. This means its compiler predetermines where all calculations will occur. This deterministic nature is crucial for efficient AI inference, as it minimizes wasted computation. Grock's approach, which emphasizes SRAM with high bandwidth over massive compute, proved more efficient for inference tasks than traditional GPUs, especially when integrated with NVIDIA's architecture. This led to an acquisition by NVIDIA for $20 billion, their largest ever, driven by the potential to significantly scale token generation within existing power and footprint constraints.
The inference inferno: exploding token consumption and demand
The transition from pre-training models to inference-time reasoning has led to an astonishing increase in token consumption. Brad Gerstner notes that Jensen Huang of NVIDIA described this shift as potentially a billion-fold increase in demand for inference capabilities. This surge, even before the advent of agents that perform actions, caused existing compute infrastructures to strain. Companies like OpenAI and Anthropic, operating with limited compute budgets (e.g., one gigawatt each), faced critical challenges in meeting this parabolic growth in demand. This underscored the urgent need for more token-efficient models and architectures.
The NVIDIA-Groq partnership: accelerating token output
A pivotal moment discussed is Sunny Madra's idea to partner Groq's deterministic, SRAM-based inference chips with NVIDIA's GPUs via NVLink Fusion. This integration allows parts of the inference process, where Groq's architecture excels, to be offloaded from NVIDIA's compute-heavy GPUs. The result is a significant increase in efficiency: the same power footprint can generate two and a half times more tokens. This collaboration was critical for NVIDIA, as it complemented their existing ecosystem and demonstrated a path to greatly enhancing inference output, a key economic driver for AI companies like OpenAI and Anthropic.
Plummeting inference costs and rising profitability
The economic landscape of AI has transformed due to rapidly decreasing inference costs. Over the past year, inference costs have fallen by approximately 90%, and by nearly 99% over two and a half years. This deflationary trend is a byproduct of extreme co-design across the entire compute 'factory,' including chips, packaging, and lithography. Concurrently, as AI capabilities become more valuable and customer willingness to pay increases, AI companies like OpenAI and Anthropic have shifted from negative to highly positive gross margins. This economic viability is crucial for sustaining the massive compute investments required.
The agent revolution and increased value delivery
Beyond generating answers, AI is moving into an 'action' phase with agents that can perform tasks like building applications, resolving customer service issues, or booking travel. While this requires an order of magnitude more tokens to execute, the value delivered to the end consumer increases by 100x. This dramatic increase in perceived value drives higher willingness to pay. Anthropic's recent substantial revenue growth illustrates how reaching a threshold of intelligent capability compels customers to adopt AI solutions, fundamentally changing the economic equation for AI services.
The future of AI: exponential growth, societal challenges, and 'bionic' adaptation
Speakers contend that we are nearing the end of an exponential curve in AI development, with capabilities accelerating faster than anticipated. While this promises an 'age of abundance,' it also presents unprecedented distribution and societal adaptation challenges. The need to manage wealth accumulation and the equitable distribution of AI's benefits is paramount. Individuals are urged to become 'bionic' by leveraging AI, as raw intellectual capability (IQ) becomes commoditized, while emotional intelligence (EQ), teamwork, and persuasion skills become more critical. The ongoing innovation, fueled by AI designing new chips and models, ensures continued progress, but necessitates active engagement with public policy and social contracts to navigate the profound changes AI will bring.
Mentioned in This Episode
●Products
●Software & Apps
●Companies
●Organizations
●People Referenced
Common Questions
Unlike traditional software, which has near-zero incremental cost of distribution, AI applications require significant compute power as more users engage, making distribution inherently more costly.
Topics
Mentioned in this video
Founder and CEO of Altimeter, a prominent investor in technology and AI, including OpenAI, Anthropic, and NVIDIA. He also started Invest America.
Co-founder of Grok and a key figure in its acquisition by NVIDIA. He has a background in serial entrepreneurship and is an expert on AI inference.
Founder of Grok and creator of the TPU at Google. He left Google to found Grok, focusing on AI inference chips.
A prominent computer scientist at Google, known for his work in distributed systems and AI. He gave a talk that inspired Jonathan Ross regarding compute limitations for speech recognition.
An investment firm founded by Brad Gerstner, managing over $15 billion. It has invested in numerous tech companies across various market cycles.
A leading AI research and deployment company. Altimeter is a significant investor, and its models like Whisper and ChatGPT were mentioned.
An AI safety and research company. Altimeter is a significant investor, and its models like Claude and Mythos were discussed.
A dominant player in AI hardware, particularly GPUs. NVIDIA acquired Grok's platform and is a key focus of the discussion regarding compute and economic scale.
A company that developed fast inference chips, competing in the AI hardware space. Brad Gerstner was on its board.
A mobile development shop co-founded by Sunny Madra and acquired by Pivotal.
Automotive company that acquired Autonomic, a smart mobility platform co-founded by Sunny Madra, where he then led their internal innovation lab, Ford X.
A smart mobility platform co-founded by Sunny Madra and acquired by Ford.
A company acquired by Grok, where Sunny Madra served as president and helped launch Grok Cloud.
Mentioned as a successful investment by Altimeter, it's a cloud-based data warehousing company.
Mentioned as a successful investment by Altimeter, it's a company focused on data streaming.
Mentioned as a successful investment by Altimeter, it's a web-based DevOps lifecycle tool.
Mentioned in the context of competing with NVIDIA's AI hardware, specifically with their Gaudi accelerators.
Mentioned as a competitor to NVIDIA in the AI hardware space, with their Instinct accelerators.
Discussed regarding its strategy for AI on devices, with concerns about privacy and the potential for its Gemini model to improve Siri.
Mentioned for its past infrastructure work in video and search, and its development of TPUs. Also mentioned in the context of Gemini models.
Mentioned in the context of Mustafa's comments on compute power relative to GPT-2 capabilities, and as a participant in Project Glasswing.
Taiwan Semiconductor Manufacturing Company, a primary player in the global semiconductor supply chain, particularly for advanced chip manufacturing.
A company co-founded by Jonathan Ross, focusing on AI inference chips. Its platform was acquired by NVIDIA.
A large language model developed by OpenAI, mentioned as an example of early AI capabilities and its evolution.
A large language model developed by Anthropic, praised for its introduction generation capabilities and discussed in the context of AI advancements.
An open-source speech recognition model from OpenAI that was made available on Grok's cloud platform.
The web browser for which vulnerabilities were found in Anthropic's Mythos model during internal testing.
A reference to a bug found in the Mythos model by Anthropic, highlighting AI's capability to find vulnerabilities in existing codebases.
An older model from OpenAI, used as a benchmark to illustrate the exponential increase in compute power and capabilities since its development.
Tensor Processing Unit, a specialized hardware accelerator developed by Google for machine learning. Jonathan Ross was the creator of the TPU at Google.
Field-Programmable Gate Array, a type of integrated circuit that can be programmed after manufacturing. Used by Jonathan Ross for the first version of the TPU.
A generation of NVIDIA GPUs that Grok's V1 chip was competitively compared against, despite being several generations older in silicon technology.
More from Stanford Online
View all 91 summaries
45 minStanford CS547 HCI Seminar | Spring 2026 | Promoting Agency in Human-AI Interaction
67 minStanford Robotics Seminar ENGR319 | Winter 2025 | Embodied Intelligence
50 minStanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Applications, AI in Life Sciences
35 minStanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Economics of Generative AI
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free