Key Moments

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work

Y CombinatorY Combinator
Science & Technology4 min read50 min video
Aug 3, 2026|5,484 views|233|6
Save to Pod

Want to know something specific about what's covered?

We've already dissected every moment. Ask and we will deliver (with timestamps).

TL;DR

Autonomous vehicles are nearly indistinguishable from human driving in safety, but the 15-year journey from demo to scaled product involves overcoming unique physical world challenges like cost of error and latency, not just digital AI issues.

Key Insights

1

A working demo of autonomous driving is at best 1% of the work; achieving the many 'nines' of reliability for a real product takes approximately 15 years and immense effort.

2

Autonomous vehicles operate with 17 times fewer serious-injury crashes than human drivers, based on over 220 million fully autonomous miles driven.

3

The four key gaps in physical AI compared to digital AI are: cost of error (lives vs. tokens), latency (milliseconds matter at speed), data (no internet equivalent), and validation (high confidence needed from day one).

4

Waymo uses a 'structure augmented end-to-end' approach, combining learned representations with materialized structure (laws of physics, rules of the road) to improve validation, training efficiency, and scaling.

5

A robust AI ecosystem for autonomous vehicles requires three components: the agent (driver), the simulator (virtual playground), and the critic (evaluator), all powered by a shared foundation model.

6

Evaluation and metrics are more critical than the model itself for physical AI, serving as the strategic advantage and the foundation for building trust and earning it through consistent, proven safety.

The 15-year leap from demo to scaled autonomous driving

Dmitri Dolgov highlights the immense difference between a functional demonstration of autonomous driving and a fully scaled, safe product. While Waymo achieved its initial demo goals, driving 100,000 miles across 10 routes without human intervention in about 1.5 years (around 2010), this represented only the first 90% of the work. The journey to a reliable product, capable of operating a large fleet and serving half a million trips weekly across 15 cities, took approximately 15 years. This extensive development period underscores that the remaining 'nines' of reliability and performance, crucial for safety-critical physical applications, require exponentially more effort than the initial demonstration. The best moments in physical AI, like a smooth maneuver that goes unnoticed by passengers, signify a task accomplished safely and smoothly, a stark contrast to the excitement of a successful demo.

Navigating the four critical gaps of physical AI

Deploying AI in the physical world presents unique challenges distinct from digital AI. Dolgov outlines four primary 'gaps': the cost of error, where mistakes can result in loss of life rather than a simple retry; latency, demanding real-time decisions within milliseconds as vehicles travel at high speeds; data scarcity, lacking the vast, pre-labeled internet of knowledge available for digital AI; and validation, requiring a very high level of confidence from day one, unlike digital AI where products can be iterated upon with user feedback. These gaps necessitate a fundamentally different approach, prioritizing safety and robustness from the outset rather than as an afterthought.

The exponential cost of reliability: counting your 'nines'

Achieving high levels of reliability and performance in physical AI is an exponential challenge. Dolgov explains that moving from 90% or 99% performance to subsequent 'nines' requires an order of magnitude more effort. For a product like an autonomous vehicle, engaging with the public, this necessitates multiple 'nines' of reliability, far beyond what's needed for digital assistants. The long tail of rare events becomes a daily reality when operating millions of miles weekly. To achieve these higher nines, fundamental shifts are required, such as building fully redundant systems and tiered fallback architectures, rather than simply doing more of the same. This reality means that while breakthroughs make demos easier, the actual product development progresses much slower, leading to frequent hype cycles where impressive demos overshadow the limited real-world product releases.

Sensor fusion: a multi-modal approach for robust perception

The choice of technology architecture is dictated by the required level of 'nines.' Waymo emphasizes a multi-modal sensing approach, using cameras, lidar, and radar. Cameras offer high resolution but are affected by lighting conditions; lidar provides direct 3D mapping but can be impacted by weather; radar excels in adverse conditions and measures velocity directly. These sensors are not backups but complement each other. By fusing data from all modalities, Waymo creates a more precise and comprehensive view of the world, crucial for detecting pedestrians in dust storms, identifying individuals in darkness, or spotting children chasing dogs, even when individual sensors might be compromised. This redundancy and complementarity are vital for safety, preventing a single obstruction like a leaf from stopping the vehicle.

Riding technological waves with unification and simplification

The rapid pace of technological advancement, particularly in AI, demands a company's ability to repeatedly adapt and integrate new breakthroughs. Waymo has rebuilt its driver around waves of innovation, from Convolutional Neural Networks (CNNs) to Transformers and now Vision-Language Models (VLMs). The true difficulty lies not just in adopting new tech for performance gains but in integrating it into a production safety-critical environment without regressions, while simultaneously reducing fragmentation and complexity. The goal is to achieve breakthrough performance through new technology while also demanding radical simplification and unification in the system architecture, enabling repeated successful adoption of innovations.

The Waymo Foundation Model: an integrated system for physical AI

At the core of Waymo's approach is its multimodal world action language model, known as the Waymo Foundation Model. This model processes diverse sensor inputs (cameras, lidar, radar) and understands the world's physics, dynamics, and social aspects. It

Common Questions

The physical world presents unique challenges: a high cost of error where mistakes can have life-threatening consequences, a critical need for low latency due to high speeds, a lack of readily available pre-labeled data like the internet, and a stringent validation gap requiring high confidence before deployment.

Topics

Mentioned in this video

More from Y Combinator

View all 617 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free