Computing Power Services: Not Everyone Can Handle the Stability Demands
At WAIC 2026, everywhere you look are signs for "AI Infra"—liquid cooling, switches, cluster software... Even resource matching platforms have changed their pitch. It looks lively, but my intuition tells me: most of these "shovel sellers" will die in the desert.
As a cross-border e-commerce operator, I've been using AI tools for product selection and copywriting for over two years. From initially tweaking APIs until my hands cramped, to now running through the entire selection-copy-advertising workflow with a set of toolchains, I have personally experienced that computing power is the "hidden trump card," but a trump card and chips are two different things. Most of the companies in the exhibition halls calling themselves "AI Infrastructure" are just repackaging old hardware businesses with new concepts. True computing power services have always been a game for the few.
Comparison: Giant Supermarkets vs. Street Corner Shops
Comparing two typical solutions clarifies the landscape.
Solution A: Cloud Vendors' "Computing Power Supermarket"
For example, elastic computing services from AWS or Alibaba Cloud allow users to spin up a GPU cluster with dozens of nodes by clicking a few times in the console. They offer standardized APIs, Kubernetes orchestration, auto-scaling, and even pre-installed images for model training frameworks.
# Assuming creating a training cluster using AWS CLI
aws sagemaker create-training-job
--training-job-name my-model
--algorithm-specification TrainingImage=...
--resource-config InstanceCount=8,InstanceType=ml.p4d.24xlarge
--output-data-config S3OutputPath=s3://my-bucket/output/
The benefit of this model is: out-of-the-box, pay-as-you-go, guaranteed stability for large-scale clusters. The downside is: expensive, and unfriendly to SMEs with non-standard scenarios.
Solution B: Small/Medium AI Infra Companies' "Custom Kitchen"
The startups in the exhibition halls often provide bundled solutions of "liquid cooling + switches + storage + scheduling software," but every version differs. They claim to help customers save 30% on costs, but actual deployment requires customers to adapt hardware, tune networks, and optimize scheduling scripts themselves.
# Example cluster scheduling configuration provided by a small company (pseudocode)
config = {
"node_type": "custom_gpu_worker",
"network_config": {
"topology": "fat_tree",
"bandwidth": "400Gbps"
},
"storage": "distributed_nfs",
"scheduler": "slurm_with_custom_plugin"
}
It looks more flexible on the surface, but the problem is: most customers simply lack the ability to handle such non-standard solutions. In the end, time costs, O&M costs, and failure risks are all passed on to the customer. I tried a company like this; the documentation read like gibberish, response to issues took three days, and eventually, I went back to cloud vendors obediently.
Core Barriers: Scale, Ecosystem, Trust
Why is computing power service a game for the few? I analyze it across three dimensions:
1. Scale Effect: Procurement prices for GPU clusters, data center electricity costs, and shared O&M teams allow big players to achieve costs 30%-50% lower than small companies. Small companies must sell at higher prices, but performance isn't necessarily better.
2. Ecosystem Lock-in: Cloud vendors have complete AI platforms (datasets, training, inference, monitoring), like SageMaker and Vertex AI. Once users adopt them, migration costs are extremely high. Small companies can only piggyback on API interfaces and cannot provide a closed loop.
3. Trust Cost: When a startup says "My cluster is cheaper than Alibaba Cloud," customers worry whether it will still exist tomorrow. Especially for enterprise clients, data security and service continuity matter more than price.
[!example] My own experience: Last year, I tried training a product recommendation model on a startup's GPU rental platform. Halfway through training, the machine went offline, and all model parameters were lost. They compensated for lost wages, but my project timeline slipped by two weeks. Since then, I only use big players' computing power services, even if they cost more.
Who Are the Real "Shovel Sellers"?
Most of those "AI Infra" companies in the exhibition halls are actually selling accessories for the "shovels"—liquid cooling cabinets, switches, server room wiring. These are extensions of traditional IT infrastructure, not AI-native computing power services. The real "shovel sellers" are the big players who can provide end-to-end services from hardware to scheduling, from training to deployment, and from monitoring to optimization.
However, there is one niche segment where a dark horse might emerge: Specialized computing power for vertical scenarios. For example, simulation platforms for autonomous driving, inference acceleration for medical imaging, real-time recommendation clusters for e-commerce. These...
Original Link: https://www.tmtpost.com/8078009.html
Physix Frontier