Open Models Are Collapsing The Cost Of AI

Y Combinator
00:56:59 Summary & quotes Report Issue
Loading transcript... Click for full transcript
About this episode Jeffrey Morgan, CEO of Ollama, outlines a strategic shift in enterprise AI adoption where open models are beco… AI summary

Jeffrey Morgan, CEO of Ollama, outlines a strategic shift in enterprise AI adoption where open models are becoming the primary driver of token consumption (80-90%) due to cost efficiency and customization, despite closed models retaining dominance for the most complex tasks. The conversation highlights the rise of coding agents and multi-agent orchestration as key use cases, the critical role of hardware acceleration from Nvidia and Apple Silicon in enabling local execution, and the necessity of solving security and safety concerns to unlock full enterprise adoption.

Key takeaways 7
  • Enterprise Adoption Trends: AT&T has already shifted 40% of its token consumption to open models, predominantly for coding agents and AI assistants. The trend is driven by a desire for control and customization, with cost being the immediate enabler.
  • Hybrid Model Strategy: The steady state for enterprises will likely see 80-90% of tokens flowing through open models (costing only 10-20% of the budget), while frontier closed models are reserved for the hardest, most complex tasks. This mirrors a law firm structure where partners (closed models) handle high-value work and associates (open models) handle volume.
  • Hardware Renaissance: Local execution is returning due to advancements in Apple Silicon (running 20B-40B parameter models effectively) and Nvidia's DGX Station/Spark (GB300 on a desk). This allows for low-latency, zero-cost local inference for straightforward tasks, creating a hybrid local-cloud execution model.
  • Security as a Gatekeeper: The primary blocker to open model adoption is security and safety. However, open models are uniquely suited for security testing (e.g., penetration testing) because they lack the restrictive safety guardrails of closed models. Solving safety governance is key to unlocking Chinese-origin models in US/EU enterprises.
  • Unbundling of the Stack: The traditional 'five-layer cake' of AI is unbundling. New opportunities exist in orchestration layers (memory, coordination, execution) that sit between the model and the application. Open source favors best-of-breed components over bundled walled gardens.
  • Rise of 'Flash' Models: Ultra-low-cost models like DeepSeek Flash are enabling a return to 'unlimited token' usage paradigms. These models are good enough for 80% of tasks and allow for high-volume orchestration where multiple cheap models solve problems more effectively than one expensive model.
  • Ollama's Origin Story: Ollama pivoted from Kubernetes security tools to local LLMs after realizing the difficulty developers faced running Llama 2 locally. The product found product-market fit quickly, growing from hobbyist Reddit users to 85% of Fortune 500 companies in roughly two years.
Notable quotes 5 AI-generated: wording and quote attribution may be wrong. Use the play link to verify.
  • “Cost is by far the largest pain point that open models can jump in and solve but you know every business has a vision of getting better control over AI and customizing it for their business and that's really their north star.”
    ▶ 0:00 Jeffrey Morgan explains the dual drivers for enterprise adoption: immediate cost savings and long-term strategic control/customization.
  • “The super majority of tokens... will be open models within a business. Call it 80 90%. That doesn't mean 89% of the the budget will go to open models. In fact... maybe you'll only pay 10 to 20% uh of of the cost towards open models, but your token most of your tokens will be going through the open models.”
    ▶ 18:19 Describing the economic steady state where open models handle volume while closed models handle complexity.
  • “If you try to get claw to like pentest your product it will just refuse to do that... Whereas there are literally obliterated uh security researcher models that you can find on hugging face that allow you to do it.”
    ▶ 7:51 Highlighting a specific use case where open models have a distinct advantage over closed frontier models due to fewer safety restrictions.
  • “Quinn 3.8 38B is now as good as Opus 4.6 for coding... incredibly exciting because you can run that on not the the the lowest memory MacBook, but the second lowest memory MacBook you can buy from the store.”
    ▶ 21:32 Illustrating how local hardware capabilities have caught up to cloud capabilities for specific high-value tasks like coding.
  • “We went from a context window of 128K to a million plus with open models. And so all that enabled this explosive growth.”
    ▶ 4:28 Explaining the technical enabler behind the surge in token usage for complex agent workflows like OpenClaw.

Chapters & Sections (25)

0:00 Open Model Enterprise Adoption Trends chapter 2
2:13 Enterprise Cost Reduction and Customization
3:37 Open Model Token Usage Growth
6:32 Open Models Security and Launch Operations chapter 2
8:28 Day Zero Model Launch Playbook
10:20 Packaging Open Models with Harnesses and Hardware
12:10 Open Model Orchestration and Unbundling chapter 2
13:58 Unbundling AI: Knowledge, Coordination, Execution
15:49 Model Unbundling and Enterprise Security
17:46 Open Closed Model Hybrid Strategy chapter 1
19:38 Hybrid Cloud and Local AI Execution
22:57 Open Models, Nvidia Hardware, and GPU Supply chapter 1
25:05 Desktop AI Hardware and GPU Supply
28:13 Open Models Enable Ultra-Low Cost AI chapter 2
30:35 Orchestration and Cost Efficiency of Open Models
33:09 Open Models Geopolitics and Security
34:38 Open Model Security and Olama Origins chapter 1
36:51 Olama's Pivot to Open Models
39:11 Olama's Pivot to Local LLMs chapter 1
42:29 Rapid Launch of Open Model Gateway
44:51 Open Model Monetization Strategy chapter 1
46:54 Open Model Monetization and Market Timing
49:39 YC Benefits and AI DevOps Lessons chapter 2
52:59 AI DevOps Lessons and LLM Non-Determinism
54:27 AI Team Structure and Inference Challenges

Transcript

Loading transcript...