Key Moments
When AI Stops Being a Project: Turning Technology into Real Value for Patients and Providers
Want to know something specific about what's covered?
We've already dissected every moment. Ask and we will deliver (with timestamps).
Key Moments
AI agents can now perform complex, multi-hour tasks, but healthcare enterprises struggle to integrate them due to siloed systems and a lack of reimagined workflows.
Key Insights
OpenAI's internal data suggests AI agents are performing tasks that take humans over 8 hours, spanning beyond engineering to finance, recruiting, and legal sectors.
In healthcare, the challenge isn't just implementing AI for tasks, but fundamentally reimagining end-to-end workflows, which requires a 'break the bone to reform it' approach.
United Health Group is consuming 10-25 billion AI tokens daily, highlighting the massive scale of AI usage and the need for robust governance, token gateways, and model management.
A recent study indicated that generalized large language models (like Claude Opus, GPT-5.2, Gemini 3.1) are outperforming specialized clinical AI models on certain benchmarks.
The industry processes 9 billion faxes annually, underscoring the vast inefficiencies and the long road ahead for AI adoption in healthcare, despite advanced model capabilities.
Moving AI from a project to the core business requires a CEO-level commitment and consistent, cross-functional (CEO, CFO, CTO, CMO) obsession with use cases, metrics, and outcomes.
AI agents are performing complex, multi-hour tasks, shifting from information retrieval to task execution.
The conversation highlights a significant evolution in AI capabilities, moving beyond simple question-answering to performing complex, multi-hour tasks. Sandeep Dadlani shared a personal anecdote about using AI to find clothing, demonstrating how AI can now handle intricate searches, compare options, and even initiate purchases. This is supported by OpenAI's internal data, which reveals that AI agents are increasingly executing tasks that would take humans over eight hours. This capability is expanding across various sectors, including engineering, finance, recruiting, and legal. The implication for healthcare is profound: AI is no longer just a tool for information retrieval but a potential executor of substantial work, pushing the boundaries of what's automated.
Healthcare's organizational structure impedes the adoption of long-form AI agentic work.
Despite the advancements in AI's task-execution capabilities, large healthcare enterprises face significant hurdles in adopting these 'long-form agentic work' applications. Sandeep Dadlani and Justin Norden both point out that current healthcare systems are not designed for such extended, AI-driven tasks. The work is often iterative and fragmented, unlike the continuous, multi-hour processes that AI agents can now handle. To truly leverage AI's potential, healthcare organizations need to fundamentally reimagine their workflows, a process described as needing to 'break the bone to reform it.' This requires moving beyond siloed AI solutions and embracing a holistic transformation of end-to-end processes, which is a significant cultural and operational challenge.
Establishing governance and reimagining workflows are critical for enterprise AI adoption.
For large enterprises like United Health Group, effectively deploying AI agents requires robust governance, guardrails, and, crucially, a reimagining of existing workflows. Sandeep Dadlani notes that while they have a sophisticated setup with 22,000 engineers and a "harness" called United AI Studio managing 117 LLM models and consuming billions of tokens daily, the biggest challenge is organizational imagination and restructuring. Simply calling APIs 'agents' or embedding AI within existing silos won't unlock its full potential. The true opportunity lies in end-to-end process redesign, enabling agents to operate with genuine agency—to reason, decide, and act across domains. This involves building secure gateways and monitoring mechanisms to manage the complexity and risks associated with widespread AI deployment, ensuring that the focus remains on value creation rather than just adopting the latest terminology.
The commoditization of models shifts value to the application layer.
A significant trend emerging is the commoditization of AI models. Papers that test models often use data that is 18 months old due to publication lags, while real-world applications are rapidly advancing. This has led to discussions around the "bitter lesson" in AI, suggesting that with enough data, time, and compute, generalized models can surpass specialized ones. While generalized LLMs like Claude Opus 4.8, GPT-5.2, and Gemini 3.1 show strong performance, the debate continues regarding their application in specific domains like healthcare. The focus is shifting from the model itself to the value created on top of it, whether through proprietary business context, specific use-case tailoring, or advanced deployment strategies. This commoditization implies that the true competitive advantage will lie in how organizations harness these models rather than in the models themselves.
Healthcare's historical inertia contrasts with AI's transformative potential.
The healthcare industry's adoption of technology has historically been cautious, with early adopters not always being the winners. This context is crucial for understanding why transformative technologies like AI often face resistance or are treated as mere projects. Processes like processing 9 billion faxes annually or relying on outdated OCR for embedded workflows highlight the deep-seated inefficiencies. While many believe AI is different and possesses the power to fundamentally change healthcare, the historical context explains why some leadership teams are hesitant. Organizations like Sandeep Dadlani's are leading by example, treating AI as a CEO-level priority with rigorous monthly reviews involving C-suite executives, emphasizing metrics, outcomes, and value. This approach is necessary to overcome inertia and realize AI's potential for order-of-magnitude improvements.
Timing and prioritizing AI initiatives require a strategic, outcome-focused approach.
Navigating the rapid pace of AI development presents a significant challenge for organizations in determining the right timing and priorities for implementation. Sandeep Dadlani shares his experience with the "100x" mantra from his time at Mars, emphasizing that large companies can move faster by getting problem-solvers—engineers and product managers—directly connected to the end-user or problem. This approach, combined with the superpowers granted by current AI tools, can accelerate innovation. Dadlani is implementing 'tiger teams' to tackle specific processes end-to-end, leveraging advanced models and encouraging rapid iteration. This aligns with Justin Norden's observation that AI is shifting from a delegated IT problem to a core CEO-level issue, requiring dedicated time and focus to drive pace and achieve tangible results, moving beyond incremental improvements to systemic transformation.
AI is becoming a fundamental 'way of working' change, not just a new tool.
The discussion strongly suggests that AI is not merely another technology to be licensed or integrated as a tool; rather, it represents a fundamental shift in how work is done. This 'way of working' change is a platform shift, akin to the transition to cloud computing, but more pervasive. Organizations that successfully embrace this transformation, moving AI from a 'project' to the 'business,' are likely to persist and thrive. The process requires significant iteration and persistence, but once the "light bulb turns on"—when teams see how AI can be directly applied to solve user needs at the moment of engagement—the results can be transformative. This involves empowering individuals closest to the problem to leverage AI's capabilities, fostering an inside-out approach where internal learnings are productized and extended across the organization to make a significant dent in healthcare.
Mentioned in This Episode
●Products
●Software & Apps
●Companies
●People Referenced
AI Agent Task Duration Estimates
Data extracted from this episode
| Task Type | Estimated Duration | Note |
|---|---|---|
| Agent task (non-engineering) | Over 8 hours | Estimated human time for completion |
| Sourcing for Father's Day gifts | Significant work | Demonstrates AI's ability to save time on complex tasks |
Optum Insight AI Studio Model Inventory
Data extracted from this episode
| Category | Count | Details |
|---|---|---|
| Total LLM models | 117 | Includes commercial and internal models |
| Internal open-weight SLM models | 20 | Built for specific purposes |
| Registered agents | 92 | Cross-domain capabilities, requiring registration for monitoring |
AI Model Performance and Cost Comparison Factors
Data extracted from this episode
| Model/Aspect | Performance | Cost | Accessibility |
|---|---|---|---|
| General LLMs (e.g., Claude Opus, GPT-5.2, Gemini 3.1) | Performing better than specialized clinical AI models (in some studies) | Potentially higher cost, token usage concerns | Generally accessible via APIs |
| Specialized Clinical AI Tools | Potentially outperformed by general LLMs in some benchmarks | Varies | Varies |
| GLM 5.2 (Open Source Chinese Model) | Achieving frontier levels, particularly in coding | Cheaper, 5-10x fewer tokens for similar outcomes compared to some others | Potentially run on own infrastructure (though requires significant hardware) |
Common Questions
AI is moving beyond basic question answering to performing complex, multi-step tasks that can take hours. This shift is seen across various industries, indicating a move towards AI agents actively 'doing the work' rather than just providing information.
Topics
Mentioned in this video
CEO of Optum Insight, discussing his experiences with AI in consumer tasks and enterprise healthcare.
CTO or lead scientist at OpenAI, who discussed the limitations of current benchmarks and the apparent lack of an upper limit for AI model performance with increased tokens.
Published a study on medical AI performance, which sparked discussion and counter-retorts from other entities.
AI used by Sandeep Dadlani to find clothing items online, demonstrating its capability in complex consumer tasks.
Mentioned as a tool with a billion users, often perceived as a Google replacement, but its capabilities are evolving towards agentic task completion.
A version of GPT mentioned as having capabilities for operationalizing tasks, highlighting the potential for AI to improve daily life.
An open-source Chinese-based AI model discussed for its cost-effectiveness and performance, particularly in coding, and its potential impact on the market.
Mentioned as a general-purpose large language model that, along with others, showed performance comparable to or better than specialized clinical AI tools in a recent study.
A product developed by Optum, aiming to solve specific workflow problems and extend capabilities across the organization.
A product developed by Optum, aimed at solving problems within healthcare workflows.
A product from Optum designed to address challenges in healthcare workflows.
More from Stanford Online
View all 136 summaries
78 minStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 17: RL Value-Based Methods
74 minStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 16: Fundamentals of RL
74 minStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 14: Intro to IL and RL
80 minStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 15: Imitation Learning
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free