Translating 'Exascale' in NVIDIA DGX SuperPOD Documentation: A Terminology Check
"Hundred-thousand-card" is not a standardized term in international technical literature. In English, one typically says "100,000 GPUs cluster" or "100k-accelerator system." But as a technical translator, I have to admit that the propagation power of the Chinese combination "Hundred-thousand-card" is far greater than "one hundred thousand accelerator cards" or "hundred-thousand-level GPU cluster." It is concise and intuitive, acting like a ruler that instantly pulls the scale of AI computing power from the "ten-thousand-card" level of previous years to the "hundred-thousand-card" level. News mentions that this cluster, named "Sugon 8000 (Dengfeng)," was built by Sugon, using entirely domestic computing power, covering full precision from FP64 to INT8. So, as a deep analysis, what I need to do is not restate this information, but starting from the obsession with translation, ask two questions: First, is the expression "Hundred-thousand-card" technically accurate? Second, compared to top-tier international solutions, where lies the true difference in the "Hundred-thousand-card Era" supported by domestic computing power?
First Question: What exactly is the "Card" in the terminology?
"Card" is an extremely imprecise umbrella term in Chinese technical documents. It can refer to GPU compute cards, AI accelerator cards, or even FPGA coprocessor cards. International literature clearly distinguishes between "accelerator" and "GPU" (Graphics Processing Unit). NVIDIA's H100 SXM is called a "GPU module," while AMD's MI300X tends to be called an "accelerator." Returning to Sugon 8000, it claims to use "entirely domestic computing power," but what is the core chip? The news doesn't specify, but according to public information, Sugon has mainly paired with Huawei Ascend, Cambricon Siyuan, and its own developed accelerator components in recent years. If these chips are not general-purpose GPUs, then the expression "Hundred-thousand-card" has semantic deviation. A more accurate translation would be "one hundred thousand AI acceleration units" or "one hundred thousand processor nodes." But Chinese media chooses "card" because it is concise and inherits the colloquial context of "graphics card." From this perspective, "Hundred-thousand-card" is a marketing term aimed at non-technical readers, not an engineering metric. As a technical translator, I tend to annotate it in documents as "Hundred-thousand-level AI Accelerator Cluster (approx. 100,000 compute units)."
Second Question: Comparison between Domestic Solutions and International Mainstream.
This is the truly interesting part. The news emphasizes "full precision coverage," from FP64 to INT8. This is a key signal. In mainstream international solutions, NVIDIA H100's FP64 performance is strictly limited (approx. 34 TFLOPS), and its core advantage lies in low-precision training and inference such as TF32, FP16, and INT8. Why do domestic clusters emphasize FP64? Because FP64 double-precision performance was once the core metric for supercomputers (scientific computing). US companies NVIDIA and AMD have long laid out in high-performance computing, but in recent years, as AI training requirements for precision gradually relaxed (mixed-precision training becoming standard), FP64 appears "expensive and inefficient." Sugon 8000 deliberately emphasizing full precision coverage, especially FP64, is likely benchmarking against traditional HPC applications (such as meteorology, oil exploration, drug molecular simulation), rather than pure AI training. This means the true positioning of this "Hundred-thousand-card" cluster is an "AI+HPC Fusion System," not a machine purely chasing large model training speed.
In comparison, NVIDIA's DGX SuperPOD or Microsoft's "Team America" clusters (such as OpenAI's hundred-thousand-card level) almost entirely adopt low-precision or mixed-precision optimizations, abandoning high-precision computing. They pursue compute per watt and throughput, not FP64 peaks. The domestic solution choosing to "grasp both hands" reflects the ambition of technological autonomy on one hand, but may also expose the problem of ecosystem fragmentation on the other—full precision support means adapting different backend compilers and drivers for different application scenarios, placing extremely high demands on software stack uniformity. In the international community, the CUDA ecosystem is highly concentrated on NVIDIA GPUs, while domestic chip vendors (Ascend, Cambricon, Tianshu Zhixin, etc.) each have different programming models and operator libraries. Whether Sugon, as the integrator, can achieve "transparent switching" of software-hardware synergy at the hundred-thousand-card scale is the key measure of the "gold content" of its "entirely domestic" claim.
My Judgment and Viewpoint
From a translation perspective, the term "Hundred-thousand-card Era" itself is not problematic; like "Ten-thousand-card Cluster," it has become a convention in the Chinese tech circle. But as a deep analysis, we need to be wary of the technical metaphors behind the vocabulary: more cards do not equal stronger computing power. International literature evaluating ultra-large-scale systems focuses more on "effective FLOPs" and "utilization rate." A hundred-thousand-card cluster, if interconnect bandwidth is insufficient and scheduling efficiency is low, may actually output less than a well-optimized fifty-thousand-card NVIDIA cluster. Sugon 8000 emphasizes "entirely domestic," but this modifier cannot be directly equated with "leading." Currently, domestic chips still have a generational gap with NVIDIA flagship products in single-card compute, memory bandwidth, and interconnect latency, especially in high-precision computing and ecosystem support. Therefore, the significance of this cluster lies more in "capability verification"—proving that China can build a hundred-thousand-card level system using domestic devices under supply chain restrictions. It is a milestone, but not an endpoint.
Core Viewpoint: The completion of the hundred-thousand-card cluster marks the starting point of China's computing infrastructure leaping from "quantitative change" to "qualitative change," but the real test lies not in how many cards are owned, but in whether these cards can work collaboratively to output effective computing power comparable to international peers. Just as translation values "Faithfulness, Expressiveness, Elegance," computing power construction values "Usable, Controllable, Scalable"—currently, we have just taken the second step.
Original Link: https://www.qbitai.com/2026/07/447891.html
Physix Frontier