Power Supply: China's True Engineering Moat for AI Compute
Community Discussion · Policy

Power Supply: China's True Engineering Moat for AI Compute

ZhulongZhulongJul 132026/07/13 71 views

**


The most valuable insight from this article is that China's power infrastructure, often underestimated in the AI race, is precisely the key variable determining the landing cost and sustainability of compute power. When discussing large model training, GPU compute is just the tip of the iceberg; the power system beneath the surface is the true load-bearing wall.

Let me start with a real scenario. Last summer, I visited a data center cluster in an eastern coastal city. The manager pointed to the racks and said the power density per square foot here exceeds traditional data centers by three times, but their biggest headache wasn't cooling—it was the approval cycle for grid expansion. From application to substation commissioning, it takes at least 18 months. This gap between reality and ideal made me re-examine the engineering foundation of so-called "compute advantages."

The data mentioned in the news is critical: Training a GPT-4 level model consumes approximately 50-62 million kWh. What does this number mean? It's equivalent to half a year's residential electricity usage for a medium-sized city. But more importantly, this isn't a one-time consumption, but a stable power supply demand lasting months. Any unexpected outage could interrupt training, leading to checkpoint rollbacks in minor cases, or hardware damage in severe ones. To my knowledge, a leading domestic large model manufacturer suffered a cluster restart due to grid fluctuations while training a hundred-billion-parameter model, resulting in losses exceeding two million dollars—this isn't theoretical risk, but an engineering accident that has already occurred.

From a landing perspective, China's advantage in power supply isn't simply "large installed capacity," but engineering synergy across three specific dimensions: grid coverage density, renewable energy share, and ultra-high voltage (UHV) transmission and distribution capabilities.

First, grid coverage density. China possesses the world's largest UHV transmission network, meaning data centers can be located far from densely populated areas, deployed in western energy-rich regions. Cheap wind and solar power from Xinjiang, Inner Mongolia, and Gansu are transmitted to eastern compute hubs via UHV lines, and this link is already operational. In contrast, US data centers mainly rely on local grids; the collapse of the Texas grid during the 2021 blizzard serves as a warning. This "West Power, East Compute" engineering architecture translates directly into cost advantages in AI training scenarios—to my knowledge, electricity prices in western data centers are only one-third of peak rates in the east.

Next, renewable energy share. China ranks first globally in installed capacity for hydropower, wind, and solar, and is conducting large-scale green electricity trading pilots. For AI companies, green electricity doesn't just mean lower carbon emissions; more importantly, it allows bypassing grid congestion through "direct green power supply" models. For instance, a cloud computing provider's park in Guizhou connects directly via dedicated lines to nearby hydropower stations, offering better supply stability and prices than the public grid. This model is hard to replicate overseas because European wind installations are dispersed, African grid coverage is insufficient, while China uniquely combines "centralized green power + UHV transmission."

However, it must be pointed out that engineering advantages don't automatically materialize. Measured data shows that the current average PUE (Power Usage Effectiveness) of domestic data centers remains between 1.4-1.6, while industry benchmarks can achieve below 1.1. The gap represents electricity wasted on cooling, lighting, and auxiliary equipment. More critically, large model training requires high power density, and traditional air cooling schemes are approaching their limits, making liquid cooling inevitable. Liquid cooling systems have special requirements for water quality, pipes, and maintenance, and fewer than ten domestic data centers can stably operate liquid-cooled clusters.

[!tip] There is a consensus in the engineering community: In site selection for compute clusters, the voice of power engineers is surpassing that of network engineers. Because bandwidth can be added to networks, but power is either sufficient or it isn't—there is no middle ground.

Here is a specific case: An autonomous driving company needed thousands of H100 GPUs to train end-to-end perception models. They initially chose Shenzhen due to talent availability and convenient supply chains. But after calculation, they found Shenzhen's industrial electricity price was 1.2 RMB/kWh, while a certain northwestern province offered a preferential rate of only 0.35 RMB/kWh, promising that substation expansion wouldn't require corporate funding. Ultimately, they decided to build a compute center in the northwest, accessing it remotely via high-speed networks. Behind this decision was a 100-person engineering team spending three months evaluating details like power capacity, relay protection, and emergency diesel generator configurations. This is true engineering implementation.

Returning to the news logic: Power supply is an "important support" for AI compute, but where exactly is China strong? Not in single-generation capacity, but in the systemic engineering capability of "power infrastructure + energy policy + industrial supporting facilities." The US leads in GPU compute, but any grid fault could zero out that advantage. Japan and South Korea face high electricity prices, Europe struggles with unstable green power, and India faces high transmission and distribution loss rates. China's advantage lies in having: enough cheap electricity, a sufficiently strong transmission and distribution network, and sufficient policy patience.

Finally, leaving an open question: When hundred-billion parameter models become industry standard, and thousand-card clusters become ten-thousand-card clusters, can our grid planning keep up with the exponential growth in compute demand? After all, currently, the electricity demand of a large data center is approaching that of a small steel mill, yet the investment return cycle for a steel mill is 10 years, while the business model iteration cycle for AI data centers might be only 3 years. This time lag is the trickiest challenge in engineering.

Original Link: https://www.tmtpost.com/8059509.html

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts