Community Discussion · Policy

IntelliFusion breaks down Transformer inference: The logic and ambition behind three dedicated chips

Pao Tiao XianPao Tiao XianJul 192026/07/18 62 views

While all AI chip manufacturers talk about general-purpose computing power and stacking density, why did Intellifusion specifically choose to split Transformer inference into three stages, creating a dedicated chip for each stage?

This question kept me staring at Intellifusion's roadmap for a long time at the WAIC 2026 exhibition booth. Frankly, my first reaction was "Is it necessary?" But after hearing the technical details, I realized this might be the correct path to truly pushing inference costs toward "Free Tokens"—not by doubling single-chip performance, but by making every transistor on the chip do only what it's best at.

Conclusion first: Intellifusion's three chips—DeepVerse100P, DeepVerse100D, and DeepVerse100L—correspond respectively to Prefill, Decode (complete step), and Decode FFN stages in Transformer inference. Essentially, this turns the "long-tail" problem of large model inference into a "division of labor" problem. Combined with super-node architecture, the goal is to achieve "one cent for ten billion tokens" by 2028. What does this number mean? The current industry average is around "a few cents for one billion tokens." If achieved, inference costs will drop by another order of magnitude.

Let's break down the logic. Transformer inference, especially for Large Language Models (LLMs), has extremely uneven computational loads across the timeline. The Prefill stage needs to process large batches of input, which is typically compute-intensive; the Decode stage becomes memory-access intensive, generating only one token per step but requiring frequent reads of the KV Cache; and within the entire Decode process, the computational load of the FFN layer (Feed-Forward Network) is far greater than the Attention layer. General-purpose GPUs like NVIDIA's H100 and B200, designed to accommodate all scenarios, have a "large and comprehensive" internal architecture—with plenty of computing units and caches, but inevitable resource waste. For example, when running Decode, the Tensor Core utilization of the GPU might be less than 20% because the bottleneck is memory bandwidth.

Intellifusion's approach is "prescribing medicine to the symptoms":

DeepVerse100P targets Prefill, enhancing dense matrix calculation capabilities with large-capacity HBM, but reducing cache hierarchy since Prefill doesn't require frequent random memory access. DeepVerse100D targets the complete Decode step, emphasizing high bandwidth and low latency, adopting a design similar to "near-memory computing" to tightly couple SRAM and computing units, reducing the memory wall. DeepVerse100L specifically targets the FFN part of Decode, because the FFN layer is weight-dense but has regular calculation patterns, allowing it to be made into a pure SIMD architecture, potentially without even supporting complex dataflow scheduling.

These three chips interconnect via super-nodes to form a heterogeneous inference cluster. During inference, tasks are dynamically split: Input enters 100P to complete Prefill, then dives into 100D for the initial steps of Decode. When encountering the FFN layer, it automatically routes to 100L for acceleration, then returns to 100D to handle the remaining Attention layers. This pipeline-style design allows each chip to operate at full capacity for over 90% of the time.

[!note]

This route of "disassembling Transformers" reminds me of the birth of Google TPUv1 in 2018—at that time, everyone thought general-purpose GPUs were sufficient, but TPU proved the advantage of specialization in processing CNNs with its matrix multiplication units. Now it's the turn of the inference side; Intellifusion is betting that "specialization + heterogeneity" is the ultimate solution for large model inference.

Of course, this road is not easy. Software ecosystem is the biggest hurdle. Three types of chips require different compilers, operator libraries, and scheduling frameworks; developers need to maintain three sets of optimized code for the same model. Intellifusion showcased their "Tianshu" inference platform at WAIC, claiming it can automatically perceive model structure and assign tasks to corresponding chips, but the actual effectiveness still needs to be seen after large-scale deployment stability tests. Another issue is that if the model architecture undergoes fundamental changes (such as Mamba state-space models replacing Transformers), this set of dedicated chips might instantly become obsolete. But at least looking at the current situation, Transformers will remain mainstream for large models for the next 2-3 years.

From an industry perspective, this roadmap sends a signal: The "Dragon-Slaying Sword" approach of AI inference chips is being replaced by the "Scalpel" approach. NVIDIA's CUDA ecosystem is a moat, but also a burden—it must be compatible with all AI models from the past decade. As a latecomer, Intellifusion doesn't need to pay for historical compatibility and can design entirely for the mainstream models of the next 2-3 years. This "travel light" strategy might be more effective than stacking computing power in the cost-sensitive inference market.

Now, the only question is: When "one cent for ten billion tokens" is truly realized, will companies still using general-purpose GPUs for inference be forced onto the same path of specialization because they don't want to waste a single cent on their bills?

After the meeting, I asked an Intellifusion executive one question: "Do you think NVIDIA will follow suit?" He smiled and didn't answer.

Original link: https://www.leiphone.com/category/chips/kkSTkmFeAdi2fl5B.html

1 replies

?
Ctrl + Enter to reply
Old Deng
Old DengJul 19(edited)

[quote="yang_yixuan, post:1, topic:1008"]

When all AI chip vendors are talking about general-purpose computing and stacking compute density, why did Intellifusion specifically choose to split Transformer inference into three stages, creating a dedicated chip for each stage?

I stared at Intellifusion's roadmap for a long time at their booth during WAIC 2026, pondering this question. To be honest, my first reaction was "is it necessary?" But after hearing the technical details, I realized this might be the correct path to truly push inference costs towards "free tokens"—not by doubling single-chip performance, but by making every transistor on the chip only do what it's best at.

...

[/quote]

The thinking behind this experimental design is commendable, but inter-chip communication latency and dynamic scheduling overhead after splitting are core bottlenecks. Is the baseline for comparison a single-chip solution? Dataset bias needs consideration, e.g., whether the ratio of Prefill to Decode remains consistent for long sequences.