Key Moments
How Anthropic builds products like Claude Code before the AI models are ready | Dianne Penn
Want to know something specific about what's covered?
We've already dissected every moment. Ask and we will deliver (with timestamps).
Key Moments
Anthropic's product team ships AI models before they're fully ready, relying on user feedback for rapid iteration and 'sweating tokens' like pixels to discover future use cases.
Key Insights
Anthropic experienced a significant inflection point with Opus 3, training a frontier model that was crucial for user reach and showcasing research, differentiating them from competitors.
Opus 45, released a year after Opus 3, was magical because it was paired with a strong product experience like Claude Code, demonstrating that frontier products are needed for frontier models to shine.
The product role at Anthropic has shifted, with 'evals are the new PRDs,' emphasizing user feedback and rigorous testing as the primary drivers of product development and value.
Spending $100,000 annually on tokens is framed as a way to experience the future of AI interaction, suggesting an 'alpha opportunity' for early adopters and builders.
Product managers are encouraged to 'sweat the tokens' as much as 'sweat the pixels,' highlighting the critical importance of hands-on, experimental interaction with AI models to drive innovation.
The 'labs' team at Anthropic operates on a thesis of identifying 'discontinuous large bets' with a strong opinion on the theme but a weekly-held opinion on the exact prototype, fostering experimentation.
From startup chaos to finding an identity
Dianne Penn joined Anthropic in 2023 when the product team comprised only five engineers, with one dedicated to the entire API business. Despite initial skepticism about Anthropic's chances against OpenAI, the company's strong culture and mission were evident from the early days. The team focused on exploring how technology could bring value to users and society, evolving from a basic chat assistant to developing capabilities like tool use. A key moment was the 'Golden Gate Bridge' experiment, where a research finding about model 'thematics' was quickly turned into a public-facing feature on their website within 24 hours. This rapid, cross-functional effort, involving engineering, product, and research, demonstrated their ability to create unique user experiences and showcase research authentically, embodying a startup pace that helped them find their identity and differentiate from competitors.
Opus 3: The first frontier model and identifying coding potential
A significant milestone for Anthropic was the training and testing of Opus 3. At a time when the company had fewer than 200 employees, it became clear they needed to create a frontier model to reach users and showcase their research. This period fostered immense trust and collaboration across inference, research, fine-tuning, and pre-training teams, even during remote winter breaks. The development of Opus 3 was not just about the model itself, but also about answering the crucial question: 'Why should somebody choose Claude?' During this phase, Dianne identified an opportunity to train models to be better at coding. While a seemingly small change from a training perspective, it allowed Anthropic to differentiate competitively, attract early enthusiasts and developers, and showcase capabilities that users hadn't thought possible at the time. This success provided foundational trust for future model development.
Opus 45 and Claude Code: The synergy of model and product
Opus 45, launched about a year after Opus 3, represented another major leap. What made Opus 45 particularly magical was the combination of a powerful model with a superior product experience, exemplified by Claude Code. The team's philosophy is that 'frontier products are needed to have frontier models.' While Claude Code provided the vehicle and user experience, Opus 45 brought the intelligence to unlock new use cases, allowing users to experience novel applications and run tasks in an agentic manner. This symbiotic relationship meant Opus 45 wouldn't have had its impact without Claude Code, and Claude Code's adoption was accelerated by Opus 45. This highlights the critical interplay between model capabilities and polished product design for user adoption and perceived value.
Navigating the exponential curve: Adaptability and first principles
The current era is characterized by an exponential curve in AI improvement, where each advancement leads to massive jumps. Navigating this pace requires adaptability and first-principles thinking. Dianne emphasizes that instead of sticking rigidly to a plan, individuals and teams must be agile, making better decisions with new information. This involves constantly reasoning about 'what's next' and 'so what,' and being willing to pivot product investments. The emerging capabilities of models often outpace predictions, leading to a positive feedback loop where new model abilities can accelerate product development. This requires the organization to have grace in bringing everyone along as the pace of change accelerates, ensuring continuous improvement and timely adaptation to the evolving AI landscape.
Emerging capabilities and the 'jagged edges' of AI
AI development, as illustrated by scaling law papers, involves not just smooth linear improvements in core capabilities but also discontinuous 'emerging capabilities.' These sudden jumps, like a model suddenly being able to perform a new calculation, can occur unpredictably. This 'jagged-edged' nature of AI development makes safety and testing crucial. Without robust evaluation (evals), these jumps might happen unnoticed. The product role, therefore, involves uncovering these emerging capabilities and understanding their implications. This also means that as models improve in one area, like agentic behavior, rough edges in other areas, such as writing, become more apparent, necessitating further focused investment and training.
Token Hacking and the value of hands-on experimentation
The idea of 'token hacking'—spending significantly on token usage to access future AI capabilities today—is presented as a strategic opportunity. Dianne frames this not just as input cost but as a means to drive experimentation and discovery. The most creative thinkers and prototypers at Anthropic spend considerable time actively using new model versions. There's no substitute for direct interaction with the technology to generate novel ideas. Internally, Anthropic fosters this through 'working in public' on Slack channels, where the entire company tests early versions, leading to emergent use cases through shared discovery and iteration. This communal approach makes experimentation less of an individual sport and more of a collective effort to unlock AI's potential.
Evals as the new PRDs: Redefining product management
The product management role at Anthropic has evolved significantly. The team now operates under the principle that 'evals are the new PRDs.' This means that rigorous user feedback and evaluation sets are the primary drivers of user value, rather than traditional Product Requirements Documents (PRDs). The process begins with understanding user pain points, which requires 'sweating the tokens' – deeply analyzing interaction transcripts to identify failure trajectories. For example, early feedback about Claude 'not following instructions' was traced to issues with JSON output. By identifying this specific pain point, the team generated eval sets to test and measure improvements, effectively creating test-driven development for product managers. This approach shortens the distance to actionability for researchers and stakeholders, allowing for continuous measurement and improvement of model capabilities.
The role of labs and 'discontinuous large bets'
Anthropic's 'labs' team is tasked with identifying and exploring 'discontinuous large bets' that might not fit into the core product roadmap. The thesis is to investigate potential 10x, 100x, or 1000x opportunities. Projects like Claude Code, Claude Skills, and Claude Design have emerged from labs. The team operates with strongly held opinions about themes but weekly-held opinions about specific prototypes, fostering a culture of experimentation. This approach allows them to pursue ambitious ideas, revisit promising concepts with newer model generations, and accelerate learning even if prototypes don't immediately ship. The success of labs is attributed to its small team structure, incredible culture, and the selection of individuals passionate about zero-to-one experimentation. This model allows for greater agility and the ability to 'see around corners' for Anthropic.
Researchers: Visionaries and iterative improvers
AI researchers at Anthropic are characterized by a dual focus: a long-term vision for what the technology can become and an immediate drive for iterative improvements. This includes ambitious, founder-like energy focused on future capabilities like AI using computers or navigating screens, alongside a practical approach to enhancing current models based on user feedback. Product managers play a critical role in bridging this gap. They translate vague user feedback, like 'Claude hallucinated,' into actionable insights for researchers by detailing the specific failure mode (e.g., tool use error, knowledge synthesis issue, alignment problem). This detailed translation, often through creating eval sets, enables researchers to address user pain points effectively and measure progress across different model versions.
The necessity of 'mind-melding' and team collaboration
In the face of rapid AI advancements, team collaboration and mutual support are paramount for sustainability and avoiding burnout. Dianne emphasizes that the work is not an individual sport, highlighting the importance of radical ownership and team collaboration. She recounts instances where team members provided extra support before launches, reviewed content, and developed demos together, demonstrating a collective effort. This 'mind-melding' or 'hive mind' approach allows individuals to take PTO with confidence that the team can manage critical tasks. This deep collaboration and low-ego, team-oriented culture are crucial for navigating high-pressure decisions and ensuring collective progress, especially as the scale of technology demands more from individuals.
The evolving product role: First principles, judgment, and hands-on experience
The skills most valued in product managers today are first-principles thinking, adaptability, and deep user empathy. Instead of pattern matching, PMs must reason through user value in the current AI context. A key shift is the emphasis on 'sweating the tokens'—deeply understanding user interactions and feedback to define actionable eval sets. For leadership, this means staying hands-on, shipping with the technology, and understanding its nuances to guide teams effectively. Managers must 'walk in the shoes' of their teams and develop this tactile understanding themselves. The joy in this work comes from exploration, experimentation, and collaboration, preferably by going deep on one or two use cases rather than spreading thinly across many. This hands-on approach ensures PMs can effectively understand and articulate AI capabilities, ultimately driving better user experiences.
Pragmatic ambition: Building forward-compatible products
The rapid pace of AI development necessitates building 'forward-compatible' products. When designing, product teams consider future model capabilities, such as what users might do with 'Claude 8.' This involves grounding ambition by defining concrete goals and then being persistent in the approach while remaining flexible on specific implementation details. The goal is not just to be ambitious, but to ensure that current efforts align with future possibilities, creating a cohesive product strategy. This requires a continuous assessment of what users will be able to achieve with more advanced AI, influencing today's development decisions to ensure long-term relevance and impact.
Safety, alignment, and the 'pushback' personality of Claude
Contrary to the intuition that safety and alignment features might limit AI capabilities, Anthropic's approach with Claude has made it more interesting and effective. The 'constitution' governing Claude's behavior encourages it to push back when appropriate, acting as a crucial thinking partner rather than a compliant assistant. This 'pushback' behavior is integrated into its core characteristics, allowing it to identify potential issues or suggest better outcomes. For example, Claude can be used to help make decisions on pricing or product strategy, not just by agreeing, but by challenging assumptions and leading to better conclusions. This proactive stance, where the AI knows 'when to push back' and when to be proactive, is seen as integral to its utility and makes it a more valuable collaborator that augments human thinking rather than merely agreeing with it.
The future of human value: Judgment, persistence, and individual voice
As AI capabilities advance, human value will likely continue to reside in areas requiring nuanced judgment, hard-earned experience, and persistence. For product leaders, this means making decisions based on a deep accumulation of experience and understanding 'what to build' and 'which bets to make.' Traits like proactivity and tenaciousness in pursuing solutions are beyond general AI capabilities. Additionally, human expertise in subject matter domains like biology or life sciences, where AI is still in its early stages, will remain critical. For children, fostering curiosity, persistence, and developing a strong, individual inner voice—being opinionated and taking a stance—are crucial for success in a future that demands unique perspectives. This emphasis on individual judgment and voice is also key for adults to avoid over-reliance on AI and maintain their own critical thinking.
The evolving role of PMs: Curiosity, detail, and user empathy
The speaker believes there will always be a need for user-centric product people who delve into the details of user needs and translate them into actionable insights for development. Even with highly capable AI models and leaning engineers, understanding 'what users are trying to accomplish' and 'bubbling that up in an actionable manner' remains a core PM function. The role demands curiosity, a hands-on approach to technology, and a willingness to be deeply involved in the details. This relentless work of understanding users and ensuring AI's impact is valuable and correct is seen as essential for Anthropic's product development and model development culture, suggesting a continued or even increased need for skilled product managers.
Finding joy and sustainability through community and focus
To combat burnout and find joy in the fast-paced AI world, the key is to avoid treating experimentation as an individual sport. Participants are encouraged to find others who share enthusiasm for AI, collaborate on use cases, and inspire each other. Focusing on one or two areas to 'go deep' rather than exploring many superficially leads to unlocking genuine value and happiness. This aligns with findings that people feel more fulfilled when AI enhances their lives significantly in specific areas. The sustainability of work in AI relies on community, mutual support, and a shared sense of purpose, allowing individuals to take breaks while trusting their team to manage ongoing critical tasks. The culture at Anthropic fosters this collaborative spirit, enabling team members to replenish and watch out for each other, making the demanding work more manageable and enjoyable.
Mentioned in This Episode
●Software & Apps
●Companies
●Organizations
●Books
●People Referenced
Common Questions
In its early days around 2021 when Diane Penn joined, Anthropic had a strong mission-driven culture with an energetic, startup-like atmosphere. Despite having only five product engineers and one engineer for their API business, the team was deeply committed to finding Anthropic's identity and delivering user value, often working bottom-up on initiatives like 'Golden Gate Claude'.
Topics
Mentioned in this video
An AI research and deployment company that Diane Penn works for, known for its Claude models. The episode discusses its early days, culture, and journey to becoming a major player in the AI space.
A competitor in the AI space, initially seen as far ahead of Anthropic. Mentioned in contrast to Anthropic's early struggles and later successes.
A fintech company offering banking services to startups and entrepreneurs, known for its product-centric approach and new conversational interface, Command.
Anthropic's flagship AI model, discussed from its early versions (Claude 2, Opus 3, Opus 45) to newer ones (Fable, Mythos). The conversation highlights its evolution, capabilities, and product integrations like Claude Code.
A product experience for coding with Claude, launched to accelerate adoption of specialized models like Opus 45.
The website for Lenny's newsletter subscribers, offering free access to AI products.
An experimental version of Claude that obsessed about the Golden Gate Bridge when a specific 'feature' was dialed up, demonstrating early interpretability research.
A sponsor of the podcast, providing APIs for enterprise features like SSO, SCIM, and audit logs for B2B SaaS companies.
A competitor AI model mentioned for its early use in coding, setting a benchmark for Claude's development in that area.
An AI model from Anthropic, whose release faced increased scrutiny and concerns about its capabilities, leading to more restrictions.
An AI model from Anthropic, which hit a new tipping point with models, facing significant scrutiny and concern upon release.
A conversational interface built into Mercury that acts as a financial operator, allowing users to query financial data and manage transactions.
President of Y Combinator, mentioned for his idea about 'token maxing' and living in the future by spending heavily on AI tokens now.
Head of Product for AI Research and Labs Teams at Anthropic and the guest on the podcast. She discusses her experiences, the company's culture, and insights into product management in AI.
CEO of Anthropic, credited with predictions about AI's capabilities in coding and the exponential curve of AI improvement.
Head of Anthropic Labs, known for his incredible vision and pushing teams to think about 10x, 100x, 1000x ideas.
Works on Anthropic Labs, mentioned as someone who has discussed the lab's work.
Author of 'Incorruptible', whose work on building great companies and having metrics around culture resonated with Diane Penn.
A book that Diane Penn used to develop a Claude skill for improving communication and management skills, particularly in difficult situations.
A book by Eric Reese about building great companies and sustaining values, which Diane Penn found insightful for team management and culture.
More from Lenny's Podcast
View all 43 summaries
73 minWhy Netflix is betting on systems thinkers—not specialists—in the AI era | Elizabeth Stone (CPTO)
97 minWhy the tech workforce is quietly splitting in two | Annual AI sentiment survey (Noam Segal)
69 minAdam Mosseri: Building Instagram for an AI world
70 minOpenAI Codex lead on the new shape of product work | Andrew Ambrosino
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free