Key Moments
What Big Tech Missed And How Startups Can Still Win
Want to know something specific about what's covered?
We've already dissected every moment. Ask and we will deliver (with timestamps).
Key Moments
World models, not LLMs, are the key to true AI common sense and robotics by learning from direct sensory experience, but they require immense compute and talent.
Key Insights
The founder's first chatbot company in 2002 was '20 years too early' as nobody understood what a chatbot was.
Selling Wit.ai to Facebook started with Mark Zuckerberg's email going to spam, which the founder initially dismissed as a scam.
Current robots are 'very dumb' despite advanced hardware because their 'brains' lack common sense, a problem world models aim to solve for open environments.
Training a foundational AI model like Ami Labs' requires billions of euros, primarily for acquiring thousands of GPUs.
Unlike LLMs trained on text about the world, world models learn directly from sensory data like video and audio, mimicking human learning from scratch.
Ami Labs raised a record-breaking 1.2 billion seed round, with the true cost being inflated external expectations rather than dilution.
The surprising genesis of a Facebook acquisition
The journey into advanced AI and entrepreneurship for Alexandre LeBrun began with an unlikely email from Mark Zuckerberg. Initially dismissing it as spam, LeBrun eventually discovered it was a genuine offer to acquire Wit.ai, his company specializing in AI and natural language processing. This experience highlighted a common startup paradox: while Y Combinator's playbook advises avoiding unsolicited contact from large corporations to prevent lowball offers, a direct, high-stakes communication from a key figure like Zuckerberg can signal a unique opportunity. The early days of Wit.ai also touched on groundbreaking concepts, as it was the first company to leverage the '.ai' domain, which at the time (early 2013) was obscure and difficult to obtain, being associated with a remote island rather than artificial intelligence.
Pioneering AI too early: Lessons from 20 years of innovation
LeBrun's entrepreneurial path is marked by a consistent tread on the cutting edge, often before the market was ready. His first venture, a chatbot company started in 2002, was so far ahead of its time that potential clients in 2002 didn't understand what a chatbot was, comparing it to live chat, which itself was a nascent concept. This resulted in a decade-long struggle to build a market. His subsequent company, Wit.ai, while not as severely early, was still ahead of the curve in conversational AI, eventually leading to its acquisition by Facebook in 2015. This pattern of being early with ambitious ideas, particularly in applied AI, has shaped his approach to founding companies, emphasizing that his strength lies in navigating nascent markets rather than solely in go-to-market strategies typical for more mature offerings.
The 'world model' bet: Moving beyond text to true understanding
Ami Labs, LeBrun's latest venture, is built on the concept of 'world models,' aiming to create AI that learns from direct sensory experience, akin to how humans and animals learn from infancy. This contrasts sharply with Large Language Models (LLMs), which are trained predominantly on text written by humans about the world. LeBrun likens LLMs to someone who has read extensively but never experienced the real world, leading to a lack of common sense and limited ability to handle novel situations. World models, by incorporating video, audio, and even tactile data (through robotics), seek to ground AI understanding in direct interaction with the physical environment. This approach is seen as essential for achieving artificial general intelligence (AGI) and enabling AI to function safely and effectively in complex, open-ended environments, particularly in fields like robotics.
Robotics and the limitations of current AI
Current robots, despite sophisticated hardware, often exhibit remarkable limitations in their intelligence. LeBrun points to videos of robots performing poorly in dynamic environments—dancing and breaking things, or engaging in karate demonstrations that could endanger bystanders. These 'dumb' robots, he argues, cannot operate safely or usefully in open environments like homes or streets. The problem is not the hardware but the 'brain.' Attempts to bridge this gap using Vision-Language Models (VLAs), which try to apply LLMs to visual tasks, are described as 'very bad hacks'—slow, inaccurate, and prohibitively expensive for continuous operation. World models offer a potential solution by providing AI with a more innate understanding of the physical world, crucial for truly autonomous and helpful robots.
The immense cost and challenge of foundational AI
Building foundational AI models requires an investment of billions, primarily driven by the insatiable demand for computing power. Ami Labs, for instance, needs thousands of GPUs, translating to significant capital expenditure. While training LLMs is also costly, LeBrun notes that training world models is comparable in expense. However, once trained, world models are expected to be more efficient for inference due to potentially fewer parameters. The difficulty isn't just securing the massive capital; it's also the challenge of acquiring and maintaining access to the necessary compute resources (GPUs), which remains a bottleneck even for well-funded startups.
Coexistence of LLMs and world models
LeBrun clarifies that the development of world models does not render LLMs obsolete. He sees LLMs as exceptionally well-suited for 'language-first' tasks and problems involving sequences of discrete symbols, such as mathematics, programming, and natural language processing. These are considered low-dimensional problems where LLMs excel. World models, on the other hand, are envisioned to be vastly superior for high-dimensional, noisy problems with long horizons that require a deep, contextual understanding of the real world, which LLMs currently lack. The future likely involves a synergy between these two types of AI, each playing to their strengths.
The contrarian bet and the pursuit of AGI
While Ami Labs' focus on world models might seem contrarian, LeBrun argues it's becoming less so. A year ago, the notion that LLMs alone wouldn't lead to Artificial General Intelligence (AGI) was met with skepticism, but now, a significant portion of the AI community agrees. The definition of AGI has shifted, and promises have been re-framed, signaling that current paths may not be leading directly to superintelligence. This growing consensus validates Ami Labs' approach, which prioritizes 'grounding' AI in real-world experience to achieve a more robust and general form of intelligence, distinct from the symbol-manipulation capabilities of LLMs.
Ambitious vision, narrow problem: Advice for founders
LeBrun advocates for a specific strategy for early-stage founders: pick an extremely narrow problem but hold an extremely large, long-term vision. He reflects on his own past mistake of trying to build a chatbot that did 'everything for everyone,' which is difficult for a small team. Instead, founders should focus on solving one specific problem for a particular industry or customer set, while clearly communicating a grander, future vision. This approach allows for credible progress in a defined area, bolstered by an ambitious underlying goal. He also draws a parallel to meta's early days, where he proposed hiring 100 data annotators for AI training; Mark Zuckerberg's response to 'hire 10,000' highlighted a cultural difference in ambition, emphasizing the importance of bold thinking even if it seems extreme. This ambition, coupled with the ability to take risks, is what allows startups to break through, a feat often harder for large, established corporations.
Mentioned in This Episode
●Software & Apps
●Companies
●Organizations
●Concepts
●People Referenced
Common Questions
LLMs learn from text written about the world, acting as a proxy, while world models learn directly from real-world sensory data like video and audio, similar to how humans or animals learn. This direct experience gives world models better common sense and the ability to handle novel situations.
Topics
Mentioned in this video
A company founded by Alex, which was one of the first to use the .ai domain and was later acquired by Facebook.
An accelerator program mentioned for its playbook on how startups should handle acquisition offers.
The company that acquired Wit AI. Later mentioned as the company where Alex worked and where he proposed an ambitious hiring plan.
Alex's first company, started in 2002, which focused on chatbots and was considered about 20 years ahead of its time.
Alex's third company, which Yann LeCun advised and invested in.
Alex's current company, focused on building foundational AI 'world models', which raised a significant seed round.
A company credited with scaling the Transformers architecture and taking risks to develop models like GPT-1, 2, 3, and ChatGPT.
Mentioned in the context of large companies that could have potentially developed Transformer-based AI but did not scale it.
The company where Alex previously worked and where early chatbot experiments were conducted. Yann LeCun also worked there.
An autonomous driving company mentioned as potentially using some flavor of world models in their technology.
Mentioned for having the theoretical idea behind Transformers in 2018 but not capitalizing on it.
A type of AI model trained on vast amounts of text data to manipulate language, good for language, math, and programming tasks.
An AI model trained directly on real-world sensory data (video, audio, touch) to gain common sense and better handle novel situations, contrasted with LLMs.
A theoretical idea related to world models that has been around for 5-10 years and is being used in some applications.
The theoretical idea behind LLMs, with papers existing since 2018 but not fully utilized until later by companies like OpenAI.
More from Y Combinator
View all 609 summaries
21 minWhy Physical AI Is the Next Platform Shift
45 minOpencode CEO: Blocked, 20X Growth in 6 Months, Building the Coding Agent for the World
31 minHow Two French Engineers In New York Built The Company That Monitors The Entire Cloud
23 minThe Model-Agnostic AI Platform Betting That No Single Lab Will Win
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free