
Token Assassin Economics: Why We Need Better 'Compute Routers,' Not More Chips
When large model companies pay several dollars per million tokens in inference costs, how many have thought about how much of that money is paying NVIDIA's ecosystem tax?
Three years ago, the clamor of the "Battle of a Hundred Models" has long faded, but a more fundamental problem has surfaced: While NVIDIA GPUs are hard to get, a large batch of domestic AI chips gather dust in data centers. This is not a capacity issue, but a systemic ecosystem mismatch—model trainers cannot migrate to non-NVIDIA hardware at low cost, while hardware vendors cannot provide sufficiently mature software stacks. The "Token Assassin" mentioned by Xia Lixue at WAIC 2026 is the ultimate manifestation of this mismatch: users pay prices far exceeding actual computing costs for compute, while the potential of domestic chips is locked in a cage of compatibility.
From "Selling Chips" to "Opening Compute Stores": A Mode Inflection Point
InfiniFlow's answer is not to build another NVIDIA, but to become a "compute middle layer." The "Open Store, Build Center" proposed by Xia Lixue's team is essentially a shift from hardware sales to compute services. This is not just a business model change, but a reconstruction of the tech stack—they attempt to establish a general runtime layer between model frameworks and underlying hardware, allowing the same model to achieve near-native inference efficiency on different chips.
This approach aligns with open-source compilers like Triton and TVM, but differs in that InfiniFlow's goal is a commercialized "compute router," not a purely academic tool. When users face the dilemma of "NVIDIA is expensive but easy, domestic cards are cheap but troublesome," the role of the middle layer is to make "troublesome" transparent, thereby breaking the pricing power of the "Token Assassin."
Technical Key: The "Glue Layer" Decoupling Models and Hardware
Achieving this goal requires solving three core problems:
- Automatic Operator Mapping: Different chips vary greatly in instruction sets and cache architectures; manually writing high-performance operators is extremely costly. InfiniFlow's approach is to build an "operator intermediate representation," letting the compiler automatically generate code adapted to different backends. This requires extensive characterization of chip micro-architecture features and is the core barrier accumulated by the team over more than a decade.
- Dynamic Compilation and Scheduling: Large model inference is dynamic; batch size, sequence length, and KV cache hit rates are constantly changing. Static compilation cannot be optimal; it must combine runtime profiles for instant tuning. This is similar to CPU out-of-order execution but on a much larger scale.
- Memory and Bandwidth Management: Domestic chips generally lag behind NVIDIA H100/B200 in memory bandwidth, requiring techniques like compute-communication overlap and pipeline parallelism to hide latency. InfiniFlow's "Center" mode can utilize multi-card, multi-node coordination, compensating for single-card weaknesses with cluster advantages.
[!info] A key judgment: In the next two years, AI chip competition will shift from "who has stronger compute" to "who is easier to integrate." The maturity of the software stack will determine the life or death of hardware vendors.
Business Logic: The Terminator of "Token Assassins"
From a purely technical perspective, InfiniFlow's solution overlaps with software stacks bundled with domestic chips like Huawei Ascend's CANN and Biren's BIRENSUITE. But Xia Lixue's differentiation lies in "neutrality"—not bound to any hardware vendor, but acting as a platform service provider aggregating multiple chip resources. For users, this means:
- Lower Migration Costs: No need to re-optimize models when switching chips; configure once on InfiniFlow's platform to switch between multiple compute pools.
- Countering Lock-in Effects: NVIDIA's CUDA ecosystem is the biggest lock-in, while the middle layer can serve as a buffer zone for "anti-lock-in." Users' data and models do not depend on specific chips; compute itself becomes a replaceable commodity.
- On-Demand Elastic Scaling: Similar to cloud computing but finer-grained—billing by Token, not by card-hour. This directly addresses the "Token Assassin" issue: users only pay for actual
Original Link: https://www.leiphone.com/category/industrynews/O4z1YxdXtRVscE16.html
Physix Frontier