The Era of GPUs Conquering All Should Come to an End
Community Discussion · Policy

The Era of GPUs Conquering All Should Come to an End

Ling XiLing XiJul 192026/07/19 53 views

Chen Ning says "Inference is productivity." It sounds like a slogan, but upon reflection, it highlights a hard truth in engineering that has been ignored for three years.

Over the past two years, everyone has been fixated on the scale of training clusters, boasting thousands or tens of thousands of cards, pushing per-card compute power higher and higher. But when it comes to deployment, we find that the bottleneck for model inference isn't floating-point peak performance, but memory bandwidth, the power wall, and actual token throughput in deployment scenarios. Vendors shouting about "100k card clusters" often get criticized by users for "slow response" and "high costs" on the inference side. Why? Because hardware requirements for training and inference follow two different logics, yet the market uses one solution—GPU—to rule them all.

[!note] Real-world Data Point: Under equivalent performance, customized inference chips can reduce unit costs to 1/5 to 1/3 of GPU solutions. This gap isn't due to a generational tech difference, but architectural mismatch.

I tried running a certain 70B model on an A100 for inference, costing about 0.2 yuan per token (calculated based on market rental rates). Using GPNPU (the architecture Chen Ning proposed) to run the same model, if mass-produced, the target cost could drop to "one cent per hundred billion tokens," i.e., 0.0001 yuan/token. This order-of-magnitude change isn't just a price cut; it completely deconstructs the business model of "inference pricing."

Let's compare the two paths:

Traditional GPU Path: General-purpose computing, flexible hardware but redundant. To handle graphics rendering and matrix multiplication, massive amounts of transistors are allocated for non-inference scenarios. During inference, most compute units sit idle, but power consumption and memory usage remain high. It's like using an aircraft carrier to deliver packages—it works, but the cost is absurd.

GPNPU Path: Designed specifically for neural network weight loading, activation computation, and sparsity handling. Unnecessary generalization is removed, leaving area for on-chip SRAM and specialized dataflow controllers. During inference, memory bandwidth utilization jumps from 30% on GPUs to over 90%, and power consumption is halved.

I fell into a pitfall: Previously, when doing edge deployment, I used a GPU to run YOLOv8. Frame rates wouldn't go up, and power consumption was high. Later, switching to an NPU doubled the frame rate and reduced power consumption by 60%. So Chen Ning's claim that "GPNPU becomes the gold standard for the inference era" isn't empty talk; it's the natural selection of engineering optimization.

But note, there's a big trap here: The GPNPU ecosystem is extremely fragmented. Everyone creates their own instruction set, spawning a bunch of LLVM-based hacked compilers. Developers adapting to different chips feel like walking through mud. Chen Ning's mention of "inference is productivity" implies a requirement for unified top-level frameworks—if inference chips can be compatible with PyTorch/TensorRT operators, the ecosystem might break through.

The attached image (from the WAIC site) directly illustrates this trend: Hardware vendors are shifting from "stacking compute power" to "compute power as a service," and the key to service is the marginal cost per token.

Actionable Advice: If you're a team implementing AI applications, you should scrutinize your inference cost structure right now. Don't blindly trust GPUs; investigate customized inference chips suited to your business scenarios (such as GPNPU, NPU, TPU). Do the math: If token costs drop by an order of magnitude, how can you reconstruct your product pricing and business model? This is the true meaning of what Chen Ning calls "inference is productivity"—not a tech slogan, but the underlying logic of business decisions.

Original Link: https://www.tmtpost.com/8070642.html

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts