Overheard in Lujiazui: Fund Managers Cut GPT-4 API Budget by Half as Internal Models Prove Sufficient
This conversation perfectly echoes the judgment made by Hugging Face CEO Clem Delangue in a TechCrunch interview a few days ago—enterprises are stopping renting AI and starting to own it. He said open-source AI is exploding, and Hugging Face has grown into the "GitHub of AI." As an analyst who looks at AI company valuations every day, this statement is worth breaking down.
Let's look at some data first. The number of models on the Hugging Face platform surpassed 2 million by the end of 2025, nearly ten times what it was two years ago. Downloads are in the trillions, but if you only stare at growth numbers, you'll miss a more critical structural change—the proportion of private repositories contributed by enterprise users jumped from 12% to 38% in the past 18 months. What does this mean? It means companies are no longer satisfied with calling a black-box API; instead, they take open-source models, feed them their own data, fine-tune them, and deploy them on their own infrastructure.
The financial logic of renting vs. owning AI is essentially no different from renting vs. buying servers. Cloud APIs have extremely high gross margins (OpenAI is around 60%-70%), but once enterprise usage crosses a certain threshold, the marginal cost advantage reverses. Take Llama 3.1 405B as an example: if an enterprise calls over 500,000 tokens per day, the per-token cost of private deployment can be squeezed down to one-fifth of the API price. More importantly, API calls are usually priced separately for input and output, whereas private deployment is primarily fixed-cost linear spending. For industries sensitive to data sovereignty and with high compliance requirements like finance, healthcare, and law, renting AI isn't even a cost issue—it's a risk issue.
Behind Delangue's mention of "no longer renting AI" lies another layer of competitive landscape deduction. Currently, AI infrastructure supply is divided into three layers: bottom-layer chips (Nvidia, AMD, custom chip players), middle-layer cloud platforms (AWS, Azure, GCP), and top-layer model APIs (OpenAI, Anthropic, Google). Hugging Face's position is subtle—it doesn't make chips, sell compute power, or directly output APIs; instead, it acts as a model distribution and collaboration platform. This is like GitHub not making operating systems, but all code runs on it. Once open-source models become the mainstream choice for enterprises, Hugging Face controls the monopolistic entry point for model distribution channels.
Of course, this narrative has competitors. Meta's Llama series takes the most aggressive open route, with weights fully public, but enterprise deployment still requires Hugging Face as a distribution and version management tool. Microsoft takes another route: embedding models into GitHub Copilot and Azure AI Studio, trying to retain renting customers through ecosystem lock-in. Google's Gemma open-source versions and Vertex AI managed services are also competing for the same market share. The core of competition isn't whether the model itself is good, but who can allow enterprises to "own" AI with the lowest friction costs.
From a valuation perspective, Hugging Face completed a new funding round in early 2025, reaching a $4.5 billion valuation. This number isn't flashy among AI star companies—OpenAI's valuation exceeds $300 billion, and Anthropic has reached $60 billion. But Hugging Face's business model is lighter, with fewer assets, and possesses network effects: more models attract more contributors, making it harder for enterprises to leave. If the trend of renting AI truly peaks and shifting to private deployment becomes the norm, then the recurring revenue growth rate of model companies dependent on API income (OpenAI, Anthropic) may see diminishing marginal returns around 2027, while Hugging Face's distribution commissions and subscription revenue elasticity will be greater.
I tend to believe that we will see a "hybrid ownership" landscape in the next two years. Enterprises will run a core open-source base model internally (like Llama 3.1 or Qwen2.5) while retaining API calls for two to three vertical scenarios. Renting remains a high-frequency, low-volume consumer good, while owning is a low-frequency, high-volume heavy asset. For investors, this means re-evaluating the quality of AI companies' revenue—is it flowing revenue from APIs, or locked-in revenue from platform subscriptions and private deployments? The latter typically commands higher multiples in valuation models due to higher stickiness and greater switching costs.
Back to that coffee shop. When checking out, I glanced at the next table, where two fund managers had already started arguing about how open Llama 4 would be. No conclusion was reached, but one thing is certain—their CTOs and CFOs no longer treat AI as a temporary rental expense. The real suspense in the market lies in who can outperform inflation in the next round of valuation models.
Original Link: https://techcrunch.com/2026/07/10/hugging-faces-ceo-on-why-companies-are-done-renting-their-ai/
Physix Frontier