Key Moments
The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Want to know something specific about what's covered?
We've already dissected every moment. Ask and we will deliver (with timestamps).
Key Moments
Cerebras' wafer-scale chips now achieve over 4,400 tokens per second, making real-time AI applications and more intelligent agents possible.
Key Insights
Cerebras' CS4 architecture doubles power and interconnect bandwidth to the wafer, and halves latency, enabling performance up to 4,400 tokens per second for models like GPT-4.
OpenAI is using Cerebras' ultra-fast inference infrastructure for critical internal use cases, including incident response and advanced research, achieving speeds 14 times faster than standard GPU performance.
The CS5 architecture, launching next year, aims for another 2x performance improvement, with projections of running medium-sized models at 10,000 tokens per second and frontier models at 5,000 tokens per second.
While NVIDIA dominates the market, OpenAI's Jalapeno chip is highlighted as a significant advancement in GPU design methodology, using an AI-first approach to build chips faster.
A key concern in the industry is the increasing dominance of Chinese labs in open-source models (estimated 95%), necessitating a national interest-level initiative for the US to maintain its hardware infrastructure competitiveness.
The future of AI compute beyond single chips lies in better integration of racks, nodes, packaging, and interconnects, with Cerebras actively pursuing DRAM stacking to enhance memory bandwidth.
Achieving unprecedented inference speeds with CS4
Cerebras' latest CS4 architecture represents a significant leap in AI inference speed, achieving over 4,400 tokens per second (TPS) on models like GPT-4. This is a nearly 2x performance increase over previous generations, made possible by a new system platform that doubles the power and interconnect bandwidth to the wafer while halving latency. The goal of making wafer-scale architecture mainstream for hyperscale data centers is being realized, addressing high-density challenges. This dramatic speed increase transforms AI capabilities, moving applications from batch processing to real-time interaction and enabling more complex, agentic workflows that require rapid iteration and reasoning.
Revolutionizing AI through ultra-fast inference
The implications of ultra-fast inference extend beyond raw speed. Cerebras CTO Sean Lie emphasizes that this capability unlocks entirely new applications and user experiences. For instance, agentic AI frameworks, which involve multiple back-and-forth interactions, become far more practical and powerful when models can process thousands of tokens per second. This enables more intelligent agents capable of deeper reasoning and more complex problem-solving. The company's transition from demonstrating the feasibility of wafer-scale technology to solving real-world problems is evident, with partnerships like the one with OpenAI highlighting the co-design of models and hardware for maximum impact.
OpenAI's strategic adoption of Cerebras infrastructure
Cerebras' partnership with OpenAI is a prime example of this new era. OpenAI is utilizing Cerebras' ultra-fast inference infrastructure to run their most capable models at speeds up to 14 times faster than traditional GPU performance. This infrastructure is strategically deployed for critical internal use cases, such as incident response teams where every second counts during service outages, and for cutting-edge research applications that benefit from enhanced reasoning capabilities. While Cerebras is committed to making ultra-fast inference broadly available, a significant portion of their current capacity is directed towards OpenAI, reflecting the transformative value they derive from this technology. This co-design spirit is seen as crucial for advancing both model intelligence and hardware speed.
Previewing CS5: Doubling down on performance
Looking ahead, Cerebras is already preparing for the next generation, CS5, built on the same modular Nexus platform as CS4. This platform is designed for multiple generations of products, ensuring a continuous performance roadmap. CS5 is projected to deliver another 2x performance improvement, aiming to enable medium-sized models like GPT-4o or Gemma to run at up to 10,000 TPS, and frontier models like GPT-5 or DeepSeek-V2 at up to 5,000 TPS. This continued push for speed underscores Cerebras' bet that the demand for ultra-fast inference is just the beginning and will continue to grow exponentially.
The significance of OpenAI's Jalapeno chip
While Cerebras focuses on wafer-scale, the discussion also touched upon OpenAI's own AI chip, Jalapeno. Sean Lie views Jalapeno as a major achievement, recognizing NVIDIA's strength in the GPU market. He highlights that OpenAI's approach to building this chip, with an 'AI-first methodology,' allowed them to develop it faster and achieve impressive results. The key takeaway for Lie isn't just the performance metrics but the innovative design methodology, which he believes is the future of the industry. Jalapeno's ability to push boundaries in both throughput and latency is seen as beneficial for the entire AI ecosystem, complementing Cerebras' offerings.
Industry consolidation and the S-RAM architecture debate
The conversation also delved into other industry players, including discussions around Grok's new chip. Lie acknowledges the growing mainstream adoption of S-RAM (SRAM) designs, a trend Cerebras has championed. However, he points out a potential limitation with non-wafer-scale S-RAM designs, suggesting that Grok's performance numbers were only shown on a smaller 30 billion parameter model. This, he infers, might be due to memory constraints on individual chips, requiring a vast number of them to hold larger models. In contrast, Cerebras' wafer-scale approach provides significantly more memory per chip, simplifying the handling of massive models.
The future of AI compute: Heterogeneity and integration
Cerebras firmly believes in a heterogeneous and disaggregated ecosystem for AI compute, where different hardware architectures serve specific workload needs. This goes beyond just prefill and decode, extending to specialized components for attention, KV cache management, and expert balancing. At the scale of hundreds of megawatts to gigawatts, disaggregation and specialization become economically viable, allowing data centers to be architected like a single, highly optimized computer. This approach involves not just compute but also memory, interconnect, and packaging, highlighting the importance of innovative integration solutions beyond the chip itself.
Navigating the US-China semiconductor supply chain dynamics
A critical geopolitical issue discussed is the growing independence of China in the semiconductor and AI model space. Lie notes that the open-source model market is now heavily dominated by Chinese labs, with an estimated 95% of high-quality open models originating from there. Concurrently, China is building its own hardware infrastructure to support these models. This situation presents a strategic challenge for the US, requiring national-level initiatives to foster innovation and maintain competitiveness in both hardware and model development. Cerebras actively supports initiatives aimed at sustaining US dominance in this critical sector.
Mentioned in This Episode
●Products
●Software & Apps
●Companies
Common Questions
The Cerebras CS4 features a new wafer-scale architecture with a modular platform, offering double the power and interconnect bandwidth, and half the latency of previous generations. It significantly boosts inference performance, enabling models like GPT-OSS to run at over 4,400 tokens per second.
Topics
Mentioned in this video
Cerebras's next-generation wafer-scale architecture, featuring a new system platform designed for hyperscale data centers with 2x power, 2x interconnect bandwidth, and half the latency of previous generations. It enables ultra-fast inference, demonstrated by running GPT-OSS at over 4,400 TPS.
A key partner of Cerebras, collaborating on co-design and utilizing Cerebras's ultra-fast inference technology for critical use cases like incident response and research applications. OpenAI also developed the Jalapeno chip.
The dominant player in the GPU market, acknowledged for its expertise and market position. While NVIDIA is pushing traditional GPU designs, the discussion suggests they may not be sufficient for ultra-fast inference regimes, prompting the industry to explore SRAM designs.
A company in the AI hardware space that has made claims about its technology but has provided limited public benchmarks or technical details, leading to skepticism about its differentiation beyond traditional GPUs.
A company noted for its work in DRAM stacking and 3D DRAM packaging, which is seen as a crucial area for innovation in AI compute integration beyond the single chip.
Mentioned in the context of its ZHBM (Zinc-based high-bandwidth memory) technology, highlighting the importance of memory advancements in AI chip development.
Mentioned as a key player in the Chinese AI hardware and model development landscape, contributing to the country's increasing independence in the semiconductor supply chain.
A model demonstrated running at over 4,400 tokens per second on the Cerebras CS4, highlighting the potential for real-time applications and more capable AI agents.
A medium-sized model that is expected to run at speeds up to 10,000 TPS on the upcoming Cerebras CS5, showcasing the advancements in inference speed.
A large language model developed in China, noted as part of the growing trend of high-quality open-source models originating from Chinese labs. Its development is supported by indigenous hardware infrastructure.
Cerebras's new modular system platform that serves as the foundation for CS4 and future generations, featuring modular power supply and server components ('backpack'). It enables 2x performance improvements and supports multiple product generations.
A chip developed by OpenAI that represents a significant leap in GPU technology, demonstrating an AI-first design methodology. It achieves impressive throughput and low latency, positioning it as a major competitor to traditional GPUs.
More from Latent Space
View all 246 summaries
84 min🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
37 min⏭️ Forward Deployed: Voice AI on what works in 2026
48 minExo: Harnesses should see their own code and logs — Alex Krentsel
96 min🔬Biology Is Turning Into Software — Matt McPartlon and Neil Patil, Chai Discovery
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free