Key Moments

Open Models Are Collapsing The Cost Of AI

Y CombinatorY Combinator
Science & Technology11 min read58 min video
Sep 4, 2026|148 views|38|3
Save to Pod

Want to know something specific about what's covered?

We've already dissected every moment. Ask and we will deliver (with timestamps).

TL;DR

Open-source AI models are rapidly closing the gap with frontier closed models in capability and are becoming the dominant force in enterprise AI, primarily driven by cost savings and coding agents.

Key Insights

1

85% of the Fortune 500 use Ollama, indicating broad enterprise adoption of open models.

2

AT&T has already shifted 40% of its token consumption to open models, predominantly for coding agents.

3

Token usage per developer on Ollama's cloud saw a 150x increase since the start of the year, largely due to coding agents and then open-claw/Hermes agent projects.

4

Chinese origin models currently dominate cloud token consumption on Ollama's platform.

5

The release cadence for open-source models is accelerating, with some models like DeepSeek Flash iterating multiple times within a single summer.

6

For enterprise, it's projected that the super majority of tokens (80-90%) will be consumed by open models, even if not the majority of the budget.

Open models are driving a significant shift in enterprise AI adoption

The AI landscape is undergoing a massive shift towards open models, particularly within enterprise settings. While businesses initially focused on fine-tuning custom models, the rapid release cycles of new models made this effort feel potentially wasted. However, the interest in customization is returning, now bolstered by the cost-effectiveness and accessibility of open-source options. Ollama, a platform for running open-source AI models, serves a substantial user base, including 85% of the Fortune 500, highlighting the widespread adoption and integration of these models into business operations. This trend is fueled by both US and Chinese-origin models, with coding agents and AI assistants being primary drivers.

Cost reduction is the primary, yet not the sole, driver for open model adoption

Cost is identified as the most significant pain point that open models can address. Businesses are increasingly looking to reduce their AI expenditure, and open models offer a compelling solution. However, the appeal extends beyond mere cost savings. Enterprises also seek greater control over their AI systems and the ability to customize models for their unique business needs. While cost is an immediate benefit, the long-term strategic goal for many companies is to leverage this cost advantage to enable deeper customization and integration of AI into their specific workflows. This dual benefit of immediate cost reduction and future customization potential makes open models a strategically attractive choice for businesses.

Coding agents and automation are leading the charge in token consumption

The exponential growth in token usage, particularly on platforms like Ollama's cloud, is predominantly driven by coding agents. These agents, powered by increasingly capable open models, can automate complex tasks, leading to a significant surge in AI interaction. Beyond developers, projects like OpenClaw and the Hermes agent have extended this automation capability to non-developers in fields such as finance, support, and marketing. This democratization of AI-powered automation, coupled with advancements like the expansion of context windows from 128K to over a million tokens, has enabled a dramatic increase in the volume of AI tasks being executed, consuming vast quantities of tokens. This surge underscores the practical impact of open models in streamlining workflows and boosting productivity across various business functions.

Chinese models are leading cloud token consumption, while US and European models are also significant

While Ollama began as a tool for running open models locally, its cloud offering has revealed significant trends in model usage. Currently, models of Chinese origin are predominantly being accessed through Ollama's cloud infrastructure. These models are utilized by businesses globally, with a notable presence in the US and Germany. This highlights a global demand for open models, regardless of their origin. The widespread adoption of these models by international businesses indicates that factors like performance, cost, and specific capabilities often outweigh geographical origin in enterprise decision-making, though security and alignment concerns remain critical considerations.

The rapid iteration cycle of open models makes custom training challenging

The pace at which new open-source models are released and updated is accelerating dramatically. For instance, the DeepSeek Flash model has seen multiple iterations within a single summer, a stark contrast to the previously slower, six-month cycle for model releases. This rapid evolution presents a challenge for businesses considering extensive custom model training, as newly fine-tuned models can quickly become outdated with subsequent releases. While tooling is improving to help teams keep pace, the sheer speed of innovation makes it difficult to stay on the cutting edge. This environment favors platforms that can quickly integrate and serve the latest models, rather than extensive in-house custom development.

AI safety and security are critical considerations, especially for Chinese-origin models

As AI capabilities advance, so do concerns around AI safety and alignment. While frontier closed models may temporarily slow down to address these issues, the open-source community continues to innovate rapidly. The release of models like GLM53, with impressive cybersecurity capabilities, presents opportunities for businesses in the security and governance space. A significant hurdle for open model adoption, particularly for Chinese-origin models, is security and safety. However, for many businesses, especially in Europe and the US, if these safety concerns can be adequately addressed, the adoption of these models becomes a very viable and exciting prospect. The ability to introspect and understand model origins is crucial for mission-critical tasks.

Open models offer unique advantages, particularly in security testing

In certain use cases, open models can offer capabilities that surpass those of frontier closed models. One key area is security testing, where specialized open-source models are available for tasks like penetration testing that are often refused by more guarded closed models. While closed models typically have built-in safety guardrails, custom-trained open models can be more amenable to aggressive security analysis. Even out-of-the-box open models often demonstrate a better understanding of user intent in security-related scenarios, allowing for more effective testing for legitimate use cases.

Ollama's role as an orchestrator for seamless model deployment

Ollama plays a crucial role in the open model ecosystem by simplifying the deployment and utilization of these models. They have developed a playbook for day-zero model launches, ensuring support across various inference engines, optimizing for speed and accuracy, and identifying appropriate use cases. This involves integrating models with harnesses (like CodeX or Open Code), ensuring reliability and capacity in the cloud, and collaborating with hardware providers like NVIDIA and Apple to optimize performance. Ollama acts as a critical layer, akin to an operating system, gluing together hardware drivers, inference layers, and application runtimes to provide a consistent and robust developer experience, tackling the complex combinatorial problem of matching harnesses to models.

The future of enterprise AI spend will heavily favor open models

The future of AI token consumption within businesses is projected to be dominated by open models, potentially accounting for 80-90% of all tokens used. While this doesn't necessarily mean an equal split in budget allocation, the sheer volume of open model usage will enable a wide array of new use cases. Frontier closed models will likely remain the go-to for the most complex and demanding tasks, where cutting-edge research and capabilities are paramount. The middle ground, however, will see a blend of open and closed models working together. This hybrid approach, akin to how organizations structure human teams with partners and associates, or how cloud computing evolved with a mix of proprietary and open-source databases, allows for cost-efficiency and broad accessibility, while reserving the most challenging problems for the most powerful, albeit more expensive, frontier models.

Local models offer cost and latency benefits, complementing cloud-based solutions

The landscape is evolving towards a hybrid model that combines both cloud-hosted and locally run AI models. Local models, particularly for less demanding tasks, offer significant advantages in terms of lower latency and reduced per-token costs, effectively leveraging upfront hardware investments. The increasing power of consumer hardware, like Apple Silicon and NVIDIA's workstation GPUs, enables the efficient running of larger models (20B to 40B parameters, sometimes up to 128B). For instance, models like Quinn 3.8 (38B) are benchmarked as equivalent to GPT-4 for coding tasks, and can run on mid-tier MacBooks. This hybrid execution model, where simpler tasks are handled locally and complex ones are routed to cloud models, promises even greater cost savings and efficiency for businesses.

The rise of 'flash models' promises unlimited token access for routine tasks

A new class of models, exemplified by DeepSeek Flash, is emerging that offers ultra-low cost per token and per task. These 'flash models' are optimized for efficiency, making them ideal for handling the majority of routine tasks. This development is key to achieving the vision of 'unlimited tokens,' reminiscent of the early ChatGPT experience where users didn't worry about token consumption. While these models may not tackle the absolute hardest problems, their speed, cost-effectiveness, and ability to be chained together through orchestration enable a vast expansion of AI use cases. This efficiency-driven approach is crucial for widespread adoption, allowing businesses to leverage AI without the constant concern of cost, effectively turning them into the workhorses of the AI ecosystem.

Geopolitical considerations and model origin are increasingly important

While many businesses prioritize where a model is run over its origin, there's a growing segment that cares deeply about model provenance. This concern extends beyond simple security to aspects like how a model 'speaks' and its cultural nuances. For mission-critical tasks, such as powering a power plant's analytics, understanding the end-to-end data lineage and model origin becomes paramount. This is particularly relevant for models like those from Neotron, which offer introspection into their training. While instances of 'booby-trapped' Chinese models haven't been widely reported, robust IT and security teams within large enterprises are accustomed to managing risks associated with open-source dependencies, a challenge not entirely new to the software world.

Ollama's entrepreneurial journey from obscure origins to widespread adoption

Ollama's journey began with a focus on developer experience, stemming from the founders' prior work on Docker Desktop. Their initial attempts at finding a problem to solve were iterative, including an application to YC with a different idea before pivoting. The company's breakthrough came with the rise of open LLMs, particularly LLaMA, which coincided with their shift to building a seamless experience for running these models locally. The name 'Ollama' itself symbolizes this focus on open models. Despite initial uncertainty and a significant period without revenue, their persistence and focus on developer needs eventually led to rapid product-market fit, transforming from a niche tool for hobbyists to a solution adopted by 85% of the Fortune 500 within a remarkably short timeframe, mirroring the speed of the PC revolution.

Monetization strategy evolved with market maturation, focusing on cloud access

Ollama's monetization strategy was deliberately patient, waiting for the open model market to mature to a point where product-market fit was clearly established. While the open-source version offered privacy-focused AI, the real opportunity for monetization emerged with the rise of use cases like coding agents that could be serviced by open models. Their cloud offering, launched after significant user adoption and market validation, allows them to capture value by providing seamless access to a wide array of open models. This approach, learned from past experiences with Docker, emphasizes aligning with customer needs and waiting for the right moment to introduce revenue streams, rather than forcing a business model prematurely.

The AI development paradigm is shifting from 'god models' to orchestrated task composition

The early vision of a single, monolithic 'god model' to handle all AI tasks is giving way to a more practical and effective approach: orchestrating multiple specialized models. While a few frontier 'god tier' models will likely continue to exist for the most complex problems, the majority of customer use cases are being solved by task composition. This involves chaining together smaller, more efficient models, which can yield better results at a lower cost, offering greater trustworthiness and repeatability. This move from a single super-model to a system of specialized agents mirrors the evolution of cloud computing, where best-of-breed solutions for individual problems have often outperformed bundled offerings. This paradigm shift opens up new avenues for startups and innovation in how AI problems are solved.

Key lessons from DevOps infrastructure are being redefined in the AI era

The AI era is forcing a re-evaluation of established principles from the DevOps and infrastructure world. Concepts like Platform-as-a-Service (PaaS), where Heroku and even early Docker operated, are seen differently; being a layer on top of something is no longer inherently a vulnerable position for startups in AI, as moving up the stack can bring you closer to the customer. Furthermore, the deterministic nature expected in traditional systems contrasts with the inherent non-determinism of LLMs, which is now seen as a feature, not a bug. Building teams also requires new approaches, as AI-driven automation reduces the need for heavy staffing in areas like customer support or cloud service delivery, presenting new challenges and opportunities in how software services are built and maintained when engineers may not fully understand every line of code.

Ollama's role in curating and simplifying the fragmented open model ecosystem

In an ecosystem flooded with diverse open models, inference technologies, and cloud services, Ollama provides a crucial curation and simplification layer. For end developers, the complexity of navigating this fragmented landscape and integrating various components can be overwhelming. Ollama, alongside initiatives like Open Router (offering a unified API for multiple model providers) and Open Code (a harness integrating with any model), abstracts away this complexity. This allows developers to focus on building their applications and software rather than wrestling with the underlying infrastructure. By bringing together models, harnesses, and cloud services into a cohesive and user-friendly experience, Ollama addresses the scarcity of readily usable AI solutions in an environment characterized by an abundance of individual components.

Common Questions

The primary drivers are cost reduction and the desire for greater control and customization of AI models for specific business needs. Open-source models offer a more accessible and adaptable solution compared to closed, proprietary models.

Topics

Mentioned in this video

Software & Apps
Cursor

A company that famously fine-tuned an AI model for its use case.

GLM

A model family known for impressive cybersecurity capabilities and rapid iteration.

Docker

A company from which some of the Olama team members came, bringing expertise in building developer experiences.

DGX Spark

A desktop hardware unit offering 128GB of unified memory, designed for running large AI models.

MLX

A project by Apple enabling LLMs to run on Mac hardware.

GitHub Copilot

An AI pair programmer that provides autocomplete suggestions, used as an example of a fast coding loop experience.

GPT Luna

An AI model that has become very price-effective, enabling widespread adoption.

Heroku

A platform as a service (PaaS) example from the previous generation of cloud computing.

Google App Engine

A platform as a service (PaaS) example from the previous generation of cloud computing.

Llama

An open-source AI model that significantly impacted the landscape and inspired the name of Olama.

Quinn

An AI model (3.8B parameters) that benchmarks show is as good as Opus 4.6 for coding.

Opus 4.6

A benchmark AI model for coding performance against which Quinn 3.8B is compared.

Gemma

AI models from DeepMind that are suitable choices for local deployment.

Neotron Ultra

A model that represents the first wave of US-based large model development.

MongoDB

A database that transitioned from developer use to enterprise adoption, used as an analogy for LLM adoption.

GPT-3

A foundational large language model, mentioned implicitly in the context of token usage and cost.

More from Y Combinator

View all 630 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free