[Tianji] Domestic Chips Moving to Training in 2026: How Hard Is This Leap?
Body:
The core argument of this article is that domestic AI chips moving from inference to training isn't just a simple performance upgrade, but a system-level reconstruction. The article breaks down the essential differences between training and inference in terms of load characteristics, compute metrics, and industrial ecosystems, pointing out the shortcomings of domestic vendors in cluster interconnects and software ecosystems, and emphasizing that TCO and stability are the hard metrics customers actually care about.
As someone who deals with compute clusters regularly, I find the description of the difference between training and inference quite accurate, especially regarding MFU and TCO as measurement dimensions—they are indeed the litmus tests for a chip's practical capability. However, one detail: the article mentions the trade-off between Scale Up and Scale Out. In actual deployment, different model architectures have vastly different weightings for these two—for example, MoE models rely more heavily on single-card memory bandwidth, while dense models depend more on interconnect bandwidth. This discrepancy leads to significant performance variations for domestic chips across different scenarios. The overall direction is correct, but for domestic chips to truly replace top-tier training clusters, the pitfalls in the software ecosystem still need to be filled gradually. I've tested some domestic cards; running inference on a single card is already decent, but stability tests for ten-thousand-card clusters currently still fail.
Original Link: [Please paste original link]
Original Article: 2026, Domestic AI Chips Crossing the Chasm: From "Inference" to "Training"
Physix Frontier