Open Models Change The Economics of AI
Y Combinator
93,630 views • 9 days ago Save 45 min 12 min read
Video Summary
The AI landscape is rapidly evolving, with open-source models emerging as a powerful force, particularly in enterprise settings. Jeffrey Morgan, CEO of Olama, highlights a significant shift towards these models, driven by their cost-effectiveness and increasing capabilities in areas like coding agents and AI assistants. While cost is a primary motivator, businesses also seek greater control and customization, with open models offering a path to achieving these goals.
Olama, a platform enabling local and cloud execution of open-source AI models, has seen explosive growth, serving millions of developers and a majority of the Fortune 500. The platform's data reveals a surge in token usage, especially with the rise of sophisticated coding agents and workflow automation tools like OpenClaw and Hermes. This trend underscores a broader movement where open models are not just cost-saving but are becoming competitive with, and in some cases superior to, frontier closed models, especially for specific use cases like security testing.
Short Highlights
- Open-source models are rapidly gaining traction in enterprise, driven by cost savings and customization needs.
- Olama, a platform for running open-source AI models, has experienced massive growth, serving millions of developers and the Fortune 500.
- Coding agents and workflow automation tools are key drivers of increased AI token usage.
- Open models are becoming increasingly competitive with closed-source frontier models, especially for specific tasks like security testing.
- The future of AI will likely involve a hybrid approach, combining open and closed-source models, as well as local and cloud deployments.
Key Details
The Rise of Open Models in Enterprise [00:01:18]
- Businesses are increasingly shifting towards open-source AI models, a trend driven by both US and Chinese model developers.
- Key use cases include coding agents and AI assistants for co-working, with specific mentions of OpenClaw and Hermes.
- Olama's platform provides insights into model usage, showing a significant increase in token flow.
"The biggest thing we're seeing is a shift to open models, especially in enterprise."
Cost as a Primary Driver [00:02:06]
- Cost is identified as the largest pain point that open models can address.
- However, businesses also aim for better control and customization of AI for their specific needs, which is their long-term goal.
- Cost savings are a short-term benefit that enables longer-term customization strategies.
"Cost is by far the largest pain point that open models can jump in and solve."
AT&T's Shift to Open Models [00:02:45]
- AT&T has reportedly shifted 40% of its token consumption to open models.
- This adoption initially favors US and European models, with Chinese models also under evaluation.
- The primary workflow driving this shift is the use of coding agents.
"I think there was a great article in the information yesterday from AT&T, and it ends up they've already shifted 40% of their token consumption to open models."
Explosive Growth in Token Usage [00:03:30]
- Token usage per developer has seen exponential growth, particularly driven by coding agents.
- The introduction of models like OpenClaw and Hermes has enabled automation of complex tasks for non-developers.
- Increased context window sizes (from 128K to over 1 million) have facilitated this growth.
"We went from a context window of 128K to a million plus with open models."
Olama's Cloud and Usage Trends [00:04:38]
- Olama's cloud service has experienced a 150x increase in token usage since the beginning of the year.
- Chinese-origin models are predominantly accessed through Olama's cloud, by businesses globally, especially in the US and Germany.
- The trend of serving out-of-the-box open models at scale only took off at the start of the current year.
"As a whole through Ollama's cloud, we saw 150X since the start of the year."
The Fine-Tuning Cycle [00:05:33]
- There was significant interest in fine-tuning custom models in early 2024, which then waned.
- This interest appears to be returning, prompting questions about whether it's a sustainable trend or another cycle.
- The rapid release of new models makes fine-tuning efforts potentially short-lived.
"Seems like it's coming back now. You have a front seat to all of it."
Accelerating Model Releases and Closing Gaps [00:06:24]
- The release cadence of open-source models is accelerating, making it harder to keep up.
- The gap between open-source and frontier closed models is closing.
- This acceleration makes custom model training more challenging, though tooling is improving.
"The release of these models, the cadence is only speeding up."
AI Safety and Open Source Growth [00:07:25]
- AI safety is becoming a major concern for frontier labs, potentially slowing their progress.
- Meanwhile, open-source models continue to improve and grow in capability.
- The release of powerful models like GLM-53 presents opportunities for security and governance startups.
"The Frontier may well slow down to figure out its alignment and containment issues."
Security and Open Models [00:08:15]
- Security and safety are significant blockers for enterprise adoption of open models.
- However, if safety concerns are addressed, Chinese-origin models are viable options for businesses.
- Open-weight models were even used by Hugging Face to detect a hack, highlighting their role in security.
"If you can solve the safety problems, by and large, adopting the Chinese Origin Model Labs is completely on the table."
Open Models for Security Testing [00:08:57]
- Open models offer advantages over closed models for security testing, such as penetration testing.
- Frontier models like Claude may refuse such requests, while specialized open models can perform them.
- Some open models are custom-trained for aggressive security testing.
"Whereas there are literally obliterated security researcher models that you can find on Hugging Face that allow you to do it."
Olama's Role in Model Launches [00:09:45]
- Model developers often contact Olama before general release to coordinate launches.
- Olama experiences significant growth spikes when new models are released.
- Operating at this scale involves a playbook for successful "day zero" model launches.
"The models are coming out faster and faster and faster."
The Challenge of Model Launches [00:10:20]
- Successful model launches require supporting the model in inference engines, ensuring speed and accuracy.
- Finding the right use cases and harnesses for developers is crucial.
- Challenges include different model architectures, tool-calling mechanics, and ensuring performance benchmarks are met.
"Every model is different. They all have different challenges and architecture changes and tool calling mechanics."
Key Components of Olama's Offering [00:11:47]
- Olama packages models with harnesses (SDKs or existing open-source options like Codex and OpenCode).
- Availability and reliability, especially in the cloud with sufficient capacity, are critical.
- Collaboration with hardware providers (Nvidia, Apple Silicon) is essential for fast and effective model execution.
"Step one is getting that right. But then you got to package it with the model, making sure that it's, you know, available, it's reliable."
The "Fire Drill" of Model Releases [00:12:58]
- Model releases often involve last-minute integration and testing, described as a "fire drill."
- Olama acts as the "glue" connecting drivers, hardware, inference layers, and application runtimes, akin to an operating system.
- This process is combinatorially complex, requiring efficient development and testing processes.
"Generally, you get the model, you're lucky if there's a, you know, the ability to access the model a few weeks in advance, a lot of this stuff comes together in the last 24 hours before the model gets released."
The Future of AI Development and Unbundling [00:14:18]
- The AI stack presents opportunities for "hidden layers" to be unbundled into specialized companies.
- Key areas include knowledge integration, coordination of agents, and execution (compute).
- This mirrors the cloud computing trend where best-of-breed services emerged instead of bundled platforms.
"It's the novelism, which is all things are just bundling or unbundling."
The New Scarcity in AI [00:15:30]
- With an abundance of open model tokens, the new scarcity lies in orchestration and agent management.
- Problems like connecting agents from point A to B are complex engineering challenges.
- The increasing capability of coding agents may reduce the need for complex, proprietary software maintenance.
"The new scarcity, the problems now are what's above the tokens, right?"
The Hybrid Model: Open vs. Closed [00:17:18]
- The future enterprise AI spend is predicted to be predominantly on open models (80-90% of tokens).
- Open models will enable more use cases due to their lower cost and abundance.
- Frontier closed models will likely be reserved for the most challenging tasks.
"The super majority of tokens, and this is our take, it will be open models within a business, call it 80, 90%."
Local vs. Cloud AI Deployment [00:20:00]
- A mix of local and cloud deployments is expected, similar to the open vs. closed model debate.
- Local models offer lower latency and cost for simpler tasks, complementing cloud models.
- New hardware is enabling powerful models (20B-120B parameters) to run effectively on personal devices.
"It's similar to the closed versus open model question. It'll be a mix in our mind."
Hybrid Execution Model [00:21:40]
- Coding agents are most effective with large cloud models, while document processing runs well locally.
- A hybrid execution model uses a router to decide whether to use local or cloud models.
- This hybrid approach further reduces costs by leveraging existing hardware for local tasks.
"And so that's where we see this hybrid execution model where some of the easier, more straightforward tasks run locally."
GPU Market Volatility [00:24:30]
- The GPU market is characterized by high volatility in prices and supply/demand.
- Accessing high-end GPUs like the B200/B300 is challenging for startups.
- Inference providers play a crucial role in pooling GPUs to meet demand.
"The prices are changing very quickly and the supply and demand volatility is, is very high there."
The "Flash Model" Revolution [00:27:15]
- New classes of models, like DeepSeek Flash, offer ultra-low cost per token and task.
- These "workhorse" models are good enough for 80% of tasks, enabling unlimited token consumption.
- Chaining these cheaper models can yield significant results through orchestration.
"It's ultra low cost per token. It's also low cost per task, which is a really important metric."
The "God Model" vs. Specialized Models [00:29:10]
- The idea of a single "God model" is giving way to orchestration of smaller, specialized models.
- Task composition with smaller models can be more repeatable, trustworthy, and cost-effective.
- While powerful "God models" will exist for specific use cases, smaller models will handle the majority of tasks.
"I think for most customer use cases, there's a, you know, level at which a model becomes good enough and then they can continue using that level of intelligence."
Geopolitics and Model Origins [00:30:30]
- Geopolitical considerations arise from model origins, with some customers prioritizing models trained in their region.
- The "how" and "where" models are run are critical for security and trust.
- Understanding model data provenance is crucial for mission-critical tasks.
"I think, you know, a lot of the geopolitical angles around this, um, start with, you know, where the, where the model's from."
The Manchurian Candidate Problem [00:31:45]
- Concerns exist about potential malicious programming in models from certain origins (the "Manchurian Candidate" problem).
- However, robust IT and security teams in large businesses are accustomed to managing open-source supply chain risks.
- Proper screening and safety checks are believed to mitigate these risks.
"For, for like critical tasks like that, how do you ensure that a Chinese model, even if it's hosted in the US, isn't basically like booby trapped to like cause problems?"
Olama's Origins and Pivot [00:33:50]
- Olama's founders previously built Docker Desktop, understanding developer experience.
- The company pivoted to focus on open-source LLMs after realizing the difficulty of running them locally.
- This pivot was driven by a desire to solve a clear customer problem with a "bias to action."
"We went from this security for Kubernetes to then like security for developers on the desktop, which is like the pivot that we've never spoken about."
The "Homebrew Computer Club" to Enterprise Speedrun [00:36:50]
- Olama's rapid adoption mirrors the speed of the PC revolution but compressed into months.
- Open models' accessibility (free to start, runnable anywhere) facilitated rapid adoption by both hobbyists and enterprise developers.
- This ease of access bypassed traditional permission hurdles in large organizations.
"That's the homebrew computer club to broad computer adoption, like speed run."
Monetization Strategy and Waiting for Maturity [00:38:20]
- Olama initially focused on developer experience without immediate monetization, similar to Docker's early days.
- The strategy was to wait for open models to mature and capture product-market fit.
- Monetization began with Olama Cloud, focusing on accessing open models for challenging problems.
"But we always felt that there was this moment where, you know, you weren't using Llama with, uh, the, the Llama models, for example, with, with tool calling right away when they came out."
Y Combinator and the Power of Peers [00:41:30]
- The founders chose Y Combinator for its ability to combat the loneliness of starting a company.
- The network and peer support, especially during the pandemic, were invaluable.
- YC's transparency and shared learnings help founders avoid common mistakes.
"Starting a company is a really lonely experience."
Lessons from Infrastructure 1.0 to AI [00:44:00]
- Traditional infrastructure paradigms (like Platform as a Service) are not always applicable to AI.
- AI's non-deterministic nature is a feature, not a bug, requiring different approaches to development and testing.
- The AI space requires new lessons in customer support, cloud service delivery, and team building.
"And in fact, going up the stack can sometimes be even better because you're closer to the customer."
Olama's Role in Curation and Simplification [00:46:00]
- Olama provides curation, simplifying a fragmented universe of models, technologies, and services.
- This allows developers to focus on building applications rather than navigating complex infrastructure.
- Services like OpenRouter and OpenCode exemplify this trend, offering unified access and harnesses.
"The end developer, to your point, they just want to build their software, right?"