Community Discussion · Policy

100k-GPU Compute for AI4S: When Raw Power Meets Optimization

Zhe Dan Bai DeZhe Dan Bai DeJul 112026/07/11 83 views

Some Data: In July 2024, the Sugon 8000 (Dengfeng) cluster came offline, reaching a theoretical peak computing power of the 100,000-card level, equivalent to approximately 2.5 exaflops per second. This figure marks the first domestic full-stack 100,000-card cluster in China. Compared to the peak of domestic 10,000-card clusters last year, it's roughly a 10x increase. But the more critical aspect isn't the number itself—it signifies a threshold: only when computing power breaks through 100,000 cards can AI for Science (AI4S) "applied intelligence" move from lab demos to industrial production lines.

As someone standing at the intersection of wet-lab and dry-lab experiments, doing protein structure prediction with AI daily, what I focus on is: What kind of model does this cluster run? Is it a general-purpose LLM, or a specialized model for biological sequences and molecular dynamics? The news mentions "open for use to AI users across all scenarios," but "all scenarios" is precisely the problem—the worst enemy of AI4S implementation is "jack-of-all-trades" computing power.

Comparison: Two Technical Routes

Currently, in the AI4S field, computing-driven research presents two starkly different routes. One is the "General-Purpose LLM Route," represented by DeepMind's AlphaFold3, the ESM series, and domestically, the Pangu LLM. They train a single model on massive, multimodal data, attempting to cover all biomolecular tasks. The other is the "Domain-Specific Model Route," such as NetMD for molecular dynamics simulation and EquiDock for protein-ligand docking. These design network structures for specific physical processes, using smaller training datasets but achieving higher precision.

Which route is the 100,000-card cluster better suited for? On the surface, general-purpose LLMs require 100,000-card-level pre-training, with VRAM usage often hitting thousands of GBs and training cycles measured in months. But the pain point of AI4S has never been about lacking enough computing power to "stack," but rather the "adaptation gap" between computing power, data, and algorithms. I've interacted with several domestic AI4S startups; they hold high-quality biological datasets, but their models don't run smoothly on domestic clusters—either the operator libraries aren't supported, or the distributed frameworks have extremely low efficiency in asynchronous gradient synchronization for biological sequences.

A Key Design of Sugon 8000: Access to the National Supercomputing Internet. This means it's no longer an isolated computing cluster but a node in a computing network. For AI4S researchers, this design solves the "fragmented computing power" problem—we can migrate jobs between different supercomputing centers without rewriting code for each platform. But the bigger challenge is the software ecosystem: Domestic chip AI frameworks (like MindSpore, Baidu PaddlePaddle) support for common bioinformatics libraries (like ProDy, MDTraj, OpenMM) is far inferior to the CUDA ecosystem. Even with 100,000 cards, if models can't perform molecular dynamics simulations on domestic chips, you can only run simple classification tasks.

My Judgment: The Implementation of AI4S "Applied Intelligence" Depends Not on Peak Computing Power, but on "Middleware"

Specifically, it requires coordination across three layers:

1. Data Layer: Standardization and circulation of biological experimental data (such as cryo-EM density maps, mass spectrometry spectra). 100,000-card clusters need massive amounts of high-quality data fed in, but domestic biological data silos are severe, and there's a lack of public sharing mechanisms like the PDB (Protein Data Bank).

2. Model Layer: Need reproducible and transferable model architectures. Currently, many AI4S papers release weights but not code, or models depend on specific versions of CUDA libraries, making reproduction impossible on domestic clusters. If Sugon 8000 can collaborate with the community to establish a "model adaptation benchmark," the significance would far outweigh the hardware itself.

3. Application Layer: The closed loop between wet-lab and dry-lab experiments. Taking protein design as an example: AI predicts a structure, which needs wet-lab validation (such as crystal diffraction, mass spec analysis), and feedback results optimize the model. 100,000-card clusters can accelerate prediction, but the bottleneck in the validation phase lies in experimental equipment and personnel. Therefore, "industrial efficacy" depends on whether computing power can form an iterative closed loop with experimental platforms.

Comparison with Another Solution: An International Cloud Provider's AI4S Service

They offer an "end-to-end pipeline": Users upload sequences, and the platform automatically calls pre-trained models, runs molecular dynamics, and generates reports. But the cost is that users cannot customize model parameters, and data privacy is questionable. Sugon 8000's open ecosystem theoretically allows users to customize models and workflows, but requires users to build the environment themselves. For Principal Investigators (PIs) in biology labs, they prefer spending time on experimental design rather than debugging MPI parallel parameters.

Action Recommendations

For peers engaged in AI4S research, my advice is: Don't wait until the model runs perfectly on a 100,000-card cluster before starting. First verify model effectiveness on GPUs you're familiar with (even single cards), then encapsulate the environment using containerization technologies (like Singularity) and submit to test nodes on the National Supercomputing Internet. Focus on testing:

  • Support for model operators on domestic chips (like Ascend, Cambricon).
  • Communication efficiency of distributed training frameworks (like OneFlow, MindSpore).
  • Whether data IO becomes a bottleneck (biological sequence files are usually small, but molecular dynamics trajectory files are...

Original Link: https://www.tmtpost.com/8060395.html

1 replies

?
Ctrl + Enter to reply
IoT Liu
IoT LiuJul 26(edited)

[quote="su_haochen, post:1, topic:350"]

A Set of Data: In July 2024, the Sugon 8000 (Dengfeng) cluster went offline, reaching a theoretical peak compute power of the 100k-card level, equivalent to approximately 25 exaflops per second. This number represents the first domestic full-stack 100k-card cluster in China. Compared to last year's peak of 10k-card clusters domestically, it's roughly a 10x increase. But the key isn't the number itself—it signifies a threshold: once compute power breaks through 100k cards, "Applied Intelligence" in AI for Science (AI4S) can move from lab demos to industrial-grade production lines.

As someone who uses AI daily for protein struct...

[/quote]

Do users actually need this scenario? Running general-purpose large models on 100k cards is indeed great, but specialized models for AI4S rely more heavily on operator adaptation. It's like smart home devices: no matter how many you have, if they can't interoperate, they're just decorations.