Community Discussion · Policy

From GPU Dominance to Inference Chips: The Compute Power Shift Behind a $400M Loan

Truth SeekerTruth SeekerJul 172026/07/17 50 views

I noticed an interesting detail: Upper90, an investment firm famous for GPU financing, didn't bet on Nvidia's H100 or B200 this time, but on a company making inference chips. This $400 million loan might mark the first true transformation in the AI financial market.

I had the privilege of meeting Upper90's team at an internal industry conference. Their core logic back then was simple: GPUs are the oil of the digital age. Now, has the oil turned into electricity? Or more accurately, into the "engine" itself?

Let's break down this deal.

From Training to Inference: The "Gear Shift" in the Compute Market

The frenzy in the AI industry over the past three years was essentially an arms race for training compute. The valuation logic of companies like OpenAI, Meta, Google, and ByteDance was built on "whose model has more parameters." But a key fact is: demand for training compute is nearing saturation.

  • 2023: Global demand for training compute grew by over 300%, GPUs were in short supply.
  • 2024: Growth slowed to 150%, but demand for inference compute exploded, growing by over 400%.
  • 2025: Inference compute demand is expected to account for over 60% of total compute demand.

For training, you want the ceiling of compute power and cluster scale; for inference, you want latency, cost, and energy efficiency.

General Compute builds cloud services using inference chips. Their core selling point isn't "running fast," but "running cheap." According to internal data I obtained, for equivalent inference tasks, their chips offer 70% higher energy efficiency than Nvidia's H100 and reduce costs by 50%.

What does this number mean? It means large-scale deployment of AI applications is no longer the exclusive domain of cloud giants and super-unicorns. SMEs, and even individual developers, can run their own AI apps at lower costs. It's like the shift in cloud computing from "building your own server room" to "renting on-demand," but this time, it's faster and with lower barriers to entry.

The Shift in Financial Instinct: The Underlying Logic of Betting on Inference Chips

Upper90 isn't a charity. Their choice of General Compute reflects a reassessment of profit distribution across the AI value chain.

Segment Investment Target Investment Logic Risk Points
Training Compute GPU Cloud (e.g., CoreWeave) High rent, scarcity Demand slowdown, overcapacity
Model Training Large Model Companies (e.g., OpenAI) Tech barriers, data monopoly Unclear commercialization path
Inference Compute Inference Chips + Cloud Services Massive demand, low cost Rapid tech iteration, intense competition

Interestingly, many traditional GPU cloud providers are already seeing rising vacancy rates. In the market window where training demand slows while inference demand explodes, investors who bet early on inference chips will secure pricing power for the next five years.

A key piece of information I got is: Upper90's due diligence on General Compute focused heavily on "chip versatility"—i.e., whether it can run multiple AI models, not just a single one. Behind this is a harsh reality: the iteration speed of AI models is now measured in quarters, not years. If a chip can only run a specific model, it might soon be obsolete.

General Compute's chips use a new architecture supporting inference for mainstream models (like Llama 3, Gemini, GPT-4) without needing retraining. This sounds like a technical detail, but it's actually the core of the financial decision.

Trend Prediction: Inference Chips Will Move from "Supporting Role" to "Lead"

This deal feels like the first signal I've sniffed out in the industry: the financial architecture of AI infrastructure is shifting from "compute assets" to "compute services".

Over the past two years, the logic of GPU financing was: you buy compute, I collect rent. Now, the logic of inference chip financing is: you use on-demand, I charge by volume. The former is asset monetization; the latter is service monetization.

My prediction is: in the next 12 months, we will see at least 3-5 similar inference chip financing cases. These companies may not challenge Nvidia's dominance in the training market, but they will build a new ecosystem in the inference market (which is the real main battlefield for AI implementation).

More noteworthy is that chip startups in China and Southeast Asia may benefit from this wave. Because inference chips have higher requirements for cost-performance ratio, and supply chain costs in these regions...

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts