News Break During LeetCode Session: AI Memory V-Die Solution Achieves 540 Token Throughput
Let's clarify what this V-Die is first. Traditional HBM (High Bandwidth Memory) stacks DRAM chips on top of logic chips, connected vertically via Through-Silicon Vias (TSV). It offers high bandwidth but puts pressure on heat dissipation because heat accumulates in the middle. The V-Die solution places DRAM chips sideways, standing them up like books on a shelf next to the chip, then uses an interconnect structure called MOSAIC to route signals and power. The direct benefit is shorter heat dissipation paths and potentially higher bandwidth. The 540 tokens/s mentioned in the news is throughput in inference scenarios, compared to about 296 tokens/s for HBM4, an increase of 82.43%. This number reminds me of a small language model inference demo I wrote—it ran so slowly on CPU it broke my heart, barely reaching single-digit tokens per second on GPU. So what does 540 tokens/s mean? It's roughly enough to support real-time responses for a medium-sized conversational model.
Thinking carefully, this improvement isn't surprising. Although HBM4 doubles bandwidth, its physical architecture dictates that heat dissipation and signal integrity are bottlenecks. The more DRAM stacking layers, the harder it is to dissipate heat. As temperature rises, frequency drops, discounting bandwidth. V-Die places DRAM sideways, turning the original vertical heat dissipation problem into a horizontal one. The chip surface area is larger, making it easier to add heatsinks or use direct air cooling. Moreover, MOSAIC interconnect uses a kind of "bridging" approach, resulting in shorter signal paths and lower latency. This reminds me of the "memory wall" question often asked in interviews—algorithms get faster, but data movement speed can't keep up. V-Die tries to bypass this wall physically, rather than just relying on stacking.
However, as a student preparing for interviews, I care more about the actual impact of this solution on AI training and inference. During training, large models frequently read and write parameters; HBM bandwidth directly determines training speed. V-Die's 540 tokens/s is inference throughput, but theoretically, training scenarios could also benefit, because increased bandwidth allows faster gradient updates and parameter synchronization. But the news only mentions inference, possibly because training scenarios are more complex, involving multi-card communication and load balancing, where improvements from a single memory solution get diluted by other bottlenecks. A classmate doing CV complained to me that when training ViT models, data loading often gets stuck on memory transfer, with GPU utilization hovering around 60%. If V-Die becomes widespread, it might push this utilization above 90%.
Another point that piques my curiosity is whether this solution competes with or complements HBM4. The news says V-Die is 82.43% higher than HBM4, but in practice, HBM4 is already in mass production, while V-Die is still in the paper stage. What the research team presented at the IEEE/JSAP conference should just be a prototype; commercialization is still far off. Also, the HBM4 ecosystem is mature, with Samsung, SK Hynix, and Micron pushing it, offering clear cost advantages. V-Die requires redesigning packaging and motherboard layouts, leading to potentially high initial costs. But in the long run, if AI chip power consumption continues to soar, heat dissipation will become a bigger constraint. HBM4's stacking limit might be around 16 layers, whereas V-Die theoretically can stack more layers without worrying about heat. So I think, in the next few years, HBM4 will continue to dominate the mainstream market, while V-Die will first test the waters in high-end custom chips or scientific research fields, such as cloud inference servers needing extreme inference throughput.
I noticed the news also mentioned MOSAIC but didn't elaborate. After looking it up, MOSAIC is an interconnect technology based on silicon interposers, integrating multiple DRAM dies and logic dies via high-density interconnects, like a "memory network." V-Die and MOSAIC might be a combined solution—V-Die handles the physical form, MOSAIC handles electrical connections. This combination makes chip design more flexible, for example, pairing DRAMs of different capacities or mixing granules of different processes. This reminds me of being asked about "heterogeneous memory systems" in interviews. Back then, I stammered through the answer, only mentioning the NUMA concept. Now looking at it, V-Die+MOSAIC is heterogeneous memory practice at the hardware level, potentially spawning new programming models in the future, such as letting developers explicitly manage memory placement locations to gain bandwidth benefits.
Finally, let's talk about trend predictions. My personal judgment is that within the next 2 to 3 years, V-Die or similar sideways DRAM solutions will appear in reference designs for data center AI accelerators.
Original Link: https://www.ithome.com/0/975/397.htm
Physix Frontier