Compute Bubble: One AI Server, 30% Working, 70% Pretending
Community Discussion · Policy

Compute Bubble: One AI Server, 30% Working, 70% Pretending

Truth SeekerTruth SeekerJul 222026/07/22 61 views

If we compare the AI industry to a gold mine, then over the past two years, everyone has been frantically digging, buying mining rigs, and hoarding them. NVIDIA's H100 and B200 are the "top-tier shovels," data centers are the "mines," and compute rental is the "miner." But recently, an unsettling fact has surfaced: The ore dug out of the mines is pitifully low in purity.

At WAIC 2026, multiple industry insiders gave a number: Less than 30%. This means that after an AI server is delivered to a user, the truly effectively utilized compute power may be less than 30% of the numbers on the spec sheet. In other words, 70% of the compute power is idling, putting on a show, turning into electricity bills and noise.

This isn't one company's biased view. Chip companies across different fields, from GPU vendors to custom AI chip firms, have unanimously shifted the conversation from "peak compute" to "effective compute." A shadow war over "compute utilization" is revealing the cracks beneath the AI industry's shiny surface.

Why is Compute Utilization So Low?

To understand this number, we first need to deconstruct the concept of "compute utilization." It's not a simple percentage, but a weighted average of complex engineering metrics. I interviewed several data center ops engineers, who provided a finer breakdown:

Stage Typical Loss Proportion
Communication Overhead in Model Training Waiting time for data transfer between GPUs 15%-25%
Memory Bandwidth Bottleneck Data movement stalls, compute units idle 20%-30%
Task Scheduling Fragmentation Multiple small tasks fail to fill a whole card 10%-20%
Software Stack Compatibility Mismatched framework, driver, and library versions 5%-10%
Power & Thermal Limits Insufficient power or thermal throttling 5%-15%

Stacking these losses, the final effective utilization falls between 20%-35%, with a median around 28%. This is the truth behind "less than 30%."

What's more painful is that this isn't even the worst case. In some small-to-medium AI companies, effective utilization may drop below 10%. Because they bought the wrong GPU models, didn't tune the software stack properly, or were just "hoarding cards" to impress investors.

Who Is Covering Up This Number?

This question touches the interests of every link in the AI supply chain. GPU vendors sell "peak compute," which is their anchor for high prices. If they told customers "you can only use 30% of what you buy," the pricing system would collapse instantly. Cloud service providers rent compute by "card-hours"; the lower the utilization, the more cards clients need to rent, meaning cloud providers make more money. AI startups need to pile up compute to boost valuations; telling investors "my compute utilization is only 20%" kills the funding story.

Thus, the three parties tacitly maintain a "numerical agreement": publicly discussing peak compute, privately complaining about utilization. Only the engineers doing the actual implementation face idle GPUs daily and know how much this bill hurts.

Compute Utilization: A Chicken-and-Egg Problem?

Some argue that low utilization is because models aren't running yet; once LLM applications become widespread, capacity will naturally fill up. Does this logic hold?

I checked training logs from several top LLM companies. The training utilization for GPT-4 level models is actually between 40%-50%, which is already very high. But during inference, utilization plummets. Because inference requests are sparse and bursty, GPUs must stay on standby constantly, unable to run continuously at full load like in training. Inference scenario compute utilization is generally below 15%.

The main scenarios for AI landing—dialogue, search, recommendation—are all inference-intensive. This means even if LLMs roll out comprehensively, overall compute utilization won't significantly improve. It might even decrease further due to increased "pre-provisioned compute" (idle resources deployed in advance to handle burst traffic).

Who Is Actually Solving the Utilization Problem?

Several undercurrents are emerging in the industry. First, software optimization companies are rising. Startups focusing on operator fusion and memory acceleration claim they can boost utilization by 15%-20%. Second, heterogeneous computing solutions are being re-emphasized. Mixing CPU, GPU, NPU, and FPGA allows different tasks to run on the most suitable hardware. Third, compute scheduling platforms are shifting from "selling cards" to "selling services," billing based on actual completed computation rather than card-hours.

But the most noteworthy action comes from the open-source community. Top AI labs are publishing training logs and efficiency tools, trying to make "utilization" a transparent metric. Once transparency is achieved, the tide will recede, and it will be obvious who is swimming naked.

Trend Prediction: By 2027, Compute Utilization Will Become a Public KPI

My judgment is: Within the next 18 months, compute utilization will become a core competitive metric in the AI industry, just like "server utilization" in cloud computing. At that point, GPU vendors will release evaluation tools for "actual compute utilization," cloud providers will write "utilization guarantees" into SLAs, and AI companies will proactively showcase "effective compute ratio" in fundraising roadshows.

  • Companies piling up compute in warehouses as decoration will be the first to be eliminated by the market.
  • Teams that can pull utilization from 30% to 50% through software optimization will earn excess returns.
  • Ultimately, the valuation logic of the AI industry will shift from "how much compute do you have" to "how much compute do you use," and from "compute scale" to "compute efficiency."

This breath needs to be taken sooner or later. Before that, many fake mining rigs pretending to dig will be carried out of the mine.

Original link: https://www.tmtpost.com/8074363.html

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts