Zhipu's 1GW Domestic Chip Data Center: A Forced 'Long March' for Compute Power
1 Gigawatt. What does this number mean? Based on current mainstream AI training cluster power consumption levels, 1 GW is roughly equivalent to the electricity demand of ten 10,000-card H100 clusters. And Zhipu has just built this data center entirely using domestic chips—not NVIDIA, not AMD, but Huawei Ascend, Cambricon, Hygon, and others. Insiders say it has already begun partial operations.
This news is worth paying attention to. Not because of the grand narrative of "domestic substitution," but because Zhipu has stepped onto a delicate timing: US export controls on chips to China are tightening layer by layer, locking down NVIDIA's H100 and A100, with even the custom H800 becoming embargoed. For Chinese LLM companies to move forward, there are only two paths—either live off existing compute pools or dive deep into domestic chips. Zhipu chose the latter, and went all-in with 1 GW right away.
Behind the 1 GW figure lies a core contradiction
Compute demand is growing exponentially, but can domestic chip supply keep up? This isn't a simple yes/no issue. Let's look at some technical details:
Comparison Item | NVIDIA H100 SXM | Huawei Ascend 910B | Cambricon MLU590
--------------------|-----------------|--------------------|------------------
FP16 Compute | 1979 TFLOPS | 640 TFLOPS | 512 TFLOPS
Interconnect Bandwidth| 900 GB/s NVLink | 400 GB/s HCCS | 200 GB/s Proprietary
Memory Capacity | 80 GB HBM3 | 64 GB HBM2e | 64 GB HBM2e
Power per Card | 700W | 310W | 350W
On paper, domestic chips lag 1-2 generations in single-card compute, with larger gaps in memory bandwidth and interconnect speed. So why dare Zhipu use them? The logic is: Large model training can leverage "scale for efficiency," using more cards and smarter parallel strategies to compensate for single-card deficits. But the prerequisite is stable inter-chip connectivity and a mature software stack.
And this is precisely the "Achilles' heel" of domestic chips.
Domestic Chips: Strengths and Weaknesses
Last year, I interviewed several domestic chip manufacturers. A consensus emerged: Hardware specs are catching up quickly, but the software ecosystem gap is at least 3 years. CUDA's dominance isn't just about compute power; it's about decades of accumulated libraries, frameworks, and toolchains. Domestic chip compilers, operator libraries, and communication libraries are still in a "catch-up" phase. Especially in large-scale distributed training scenarios, stability issues amplify exponentially.
Zhipu isn't treating domestic chips as "workhorses" but as "partners" to fine-tune. This requires massive engineering investment. As far as I know, Zhipu's tech team has been deeply involved in optimizing Ascend's software stack since last year, even participating in defining interfaces for Huawei's CANN (Heterogeneous Computing Architecture). This "co-research" model allowed Zhipu to run its GLM series models on domestic chips, achieving performance close to 80% of the H800.
[!note] Another implication of this news is: Zhipu is using actual deployment to force the domestic chip industry chain to "break out." A 1 GW scale means domestic chips must move from "samples" to "mass production," from "lab" to "production line." This is a huge test for Huawei, Cambricon, and others regarding yield rates, delivery capabilities, and after-sales support.
Why Zhipu and not other players?
Among domestic LLM companies, Zhipu's "background" is somewhat special. It originated from a Tsinghua University team, with "Zhipu" in its name, backed by academic support from the Beijing Academy of Artificial Intelligence (BAAI). Compared to internet giants like ByteDance and Alibaba, Zhipu is more of a "tech faction": focusing on model architecture innovation rather than application landing. But it's exactly this "tech faction" temperament that makes it daring enough to "bet" on domestic chips.
Note that if the compute foundation is unstable, model training could interrupt at any time, with a single checkpoint restart potentially costing millions of kWh in electricity. Zhipu's 1 GW data center is essentially betting that the TCO (Total Cost of Ownership) of domestic chips can beat time. Currently, electricity cost is secondary; the real hidden costs are the "tuition fees" for software adaptation, ops troubleshooting, and architectural iteration.
Actionable Advice: Watch Three Clues
As an industry observer, I won't simply say "the prospects for domestic substitution are broad." I suggest you focus on these three specific directions:
- Software Open Source Strategy of Domestic Chip Makers: See if Ascend and Cambricon are opening more source code for operator and communication libraries, and if there is community contribution. This is a key metric for measuring ecosystem maturity.
- Zhipu's Future Model Training Efficiency Reports: If Zhipu publishes MFU (Model FLOPs Utilization) data for its domestic chip clusters, it will be an extremely valuable benchmark. Currently, H100 MFU is generally 50%-60%. If domestic chips can achieve over 40%, it proves they are "usable."
- Policy Signals: Will There Be "Compute Voucher" Subsidies? Similar to new energy subsidies, if the government provides electricity price or procurement subsidies for domestic chip compute, it will accelerate ecosystem formation.
Finally, this image might explain the situation better: ![](https://bbs-physixfrontier-com-data.oss-cn-hongkong.aliyuncs.com/collector/article
Original link: https://www.ithome.com/0/979/301.htm
Physix Frontier