Community Discussion · Policy

[Tianji] Domestic Chips Moving to Training in 2026: How Hard Is This Leap?

Tian JiTian JiJul 82026/07/08 125 views

Body:

The core argument of this article is that domestic AI chips moving from inference to training isn't just a simple performance upgrade, but a system-level reconstruction. The article breaks down the essential differences between training and inference in terms of load characteristics, compute metrics, and industrial ecosystems, pointing out the shortcomings of domestic vendors in cluster interconnects and software ecosystems, and emphasizing that TCO and stability are the hard metrics customers actually care about.

As someone who deals with compute clusters regularly, I find the description of the difference between training and inference quite accurate, especially regarding MFU and TCO as measurement dimensions—they are indeed the litmus tests for a chip's practical capability. However, one detail: the article mentions the trade-off between Scale Up and Scale Out. In actual deployment, different model architectures have vastly different weightings for these two—for example, MoE models rely more heavily on single-card memory bandwidth, while dense models depend more on interconnect bandwidth. This discrepancy leads to significant performance variations for domestic chips across different scenarios. The overall direction is correct, but for domestic chips to truly replace top-tier training clusters, the pitfalls in the software ecosystem still need to be filled gradually. I've tested some domestic cards; running inference on a single card is already decent, but stability tests for ten-thousand-card clusters currently still fail.

Original Link: [Please paste original link]

Original Article: 2026, Domestic AI Chips Crossing the Chasm: From "Inference" to "Training"

1 replies

?
Ctrl + Enter to reply
Old Ye from BCG
Old Ye from BCGJul 27(edited)

[quote="tianji, post:1, topic:119"]

Body:

The core viewpoint of this article is that domestic AI chips moving from inference to training is not a simple performance upgrade, but a system-level reconstruction. The article breaks down the essential differences between training and inference in terms of load characteristics, computing power metrics, and industrial ecosystem, pointing out the shortcomings of domestic vendors in cluster interconnection and software ecosystem, and emphasizing that TCO and stability are the hard metrics customers truly care about.

As someone who deals with computing clusters regularly, I think the description of the difference between training and inference in the article is spot-on, especially the MFU and TCO dimensions, which are indeed litmus tests for chip combat capability. But there's a detail…

[/quote]

Looking at it from three dimensions: hardware specs are just the entry ticket; the real bottleneck lies in cluster stability and software stack maturity. The MoE vs. dense model difference you mentioned perfectly illustrates that domestic chips can currently only "pick scenarios," still far from generalized replacement. Suggest advancing in stages: first tackle benchmark cases in single scenarios, then gradually converge the ecosystem gap.