About this episodeJeffrey Morgan, CEO of Ollama, outlines a strategic shift in enterprise AI adoption where open models are beco…AI summary
Jeffrey Morgan, CEO of Ollama, outlines a strategic shift in enterprise AI adoption where open models are becoming the primary driver of token consumption (80-90%) due to cost efficiency and customization, despite closed models retaining dominance for the most complex tasks. The conversation highlights the rise of coding agents and multi-agent orchestration as key use cases, the critical role of hardware acceleration from Nvidia and Apple Silicon in enabling local execution, and the necessity of solving security and safety concerns to unlock full enterprise adoption.
Key takeaways 7
Enterprise Adoption Trends: AT&T has already shifted 40% of its token consumption to open models, predominantly for coding agents and AI assistants. The trend is driven by a desire for control and customization, with cost being the immediate enabler.
Hybrid Model Strategy: The steady state for enterprises will likely see 80-90% of tokens flowing through open models (costing only 10-20% of the budget), while frontier closed models are reserved for the hardest, most complex tasks. This mirrors a law firm structure where partners (closed models) handle high-value work and associates (open models) handle volume.
Hardware Renaissance: Local execution is returning due to advancements in Apple Silicon (running 20B-40B parameter models effectively) and Nvidia's DGX Station/Spark (GB300 on a desk). This allows for low-latency, zero-cost local inference for straightforward tasks, creating a hybrid local-cloud execution model.
Security as a Gatekeeper: The primary blocker to open model adoption is security and safety. However, open models are uniquely suited for security testing (e.g., penetration testing) because they lack the restrictive safety guardrails of closed models. Solving safety governance is key to unlocking Chinese-origin models in US/EU enterprises.
Unbundling of the Stack: The traditional 'five-layer cake' of AI is unbundling. New opportunities exist in orchestration layers (memory, coordination, execution) that sit between the model and the application. Open source favors best-of-breed components over bundled walled gardens.
Rise of 'Flash' Models: Ultra-low-cost models like DeepSeek Flash are enabling a return to 'unlimited token' usage paradigms. These models are good enough for 80% of tasks and allow for high-volume orchestration where multiple cheap models solve problems more effectively than one expensive model.
Ollama's Origin Story: Ollama pivoted from Kubernetes security tools to local LLMs after realizing the difficulty developers faced running Llama 2 locally. The product found product-market fit quickly, growing from hobbyist Reddit users to 85% of Fortune 500 companies in roughly two years.
Notable quotes 5AI-generated: wording and quote attribution may be wrong. Use the play link to verify.
“Cost is by far the largest pain point that open models can jump in and solve but you know every business has a vision of getting better control over AI and customizing it for their business and that's really their north star.”
▶ 0:00Jeffrey Morgan explains the dual drivers for enterprise adoption: immediate cost savings and long-term strategic control/customization.
“The super majority of tokens... will be open models within a business. Call it 80 90%. That doesn't mean 89% of the the budget will go to open models. In fact... maybe you'll only pay 10 to 20% uh of of the cost towards open models, but your token most of your tokens will be going through the open models.”
▶ 18:19Describing the economic steady state where open models handle volume while closed models handle complexity.
“If you try to get claw to like pentest your product it will just refuse to do that... Whereas there are literally obliterated uh security researcher models that you can find on hugging face that allow you to do it.”
▶ 7:51Highlighting a specific use case where open models have a distinct advantage over closed frontier models due to fewer safety restrictions.
“Quinn 3.8 38B is now as good as Opus 4.6 for coding... incredibly exciting because you can run that on not the the the lowest memory MacBook, but the second lowest memory MacBook you can buy from the store.”
▶ 21:32Illustrating how local hardware capabilities have caught up to cloud capabilities for specific high-value tasks like coding.
“We went from a context window of 128K to a million plus with open models. And so all that enabled this explosive growth.”
▶ 4:28Explaining the technical enabler behind the surge in token usage for complex agent workflows like OpenClaw.
Chapters & Sections (25)▼
0:00Open Model Enterprise Adoption Trendschapter2
2:13Enterprise Cost Reduction and Customization
3:37Open Model Token Usage Growth
6:32Open Models Security and Launch Operationschapter2
8:28Day Zero Model Launch Playbook
10:20Packaging Open Models with Harnesses and Hardware
12:10Open Model Orchestration and Unbundlingchapter2