Community Discussion · Policy

As a complete novice, I tried Alibaba's 2.4 trillion parameter model

Is Operator Fusion Done?Is Operator Fusion Done?Aug 72026/08/07 180 views

As someone who doesn't know much about this, I tried Alibaba's Qwen3.8-Max. Wait, actually I do know a bit—I ran Qwen's open-source models on Ascend hardware for two weeks, so I have some idea about their operator efficiency. But honestly, seeing the 2.4 trillion parameter scale, my first reaction was: how do you even deploy this? It won't fit in single-machine VRAM, and multi-machine communication bandwidth becomes a bottleneck.

1 replies

?
Ctrl + Enter to reply
Long Ji
Long JiAug 7

I know way too much about the communication overhead. I tested MoE the week before last; running it on a single machine with 8 GPUs, just waiting for all-to-all was enough to make me drink a pot of tea (exhausting). Don't count on single machines. Going straight to 9 cards causes a cliff dive—in actual tests, it's slower than 8 cards, can you believe that? As for inference costs, it looks cheap now, but once you actually run it, communication costs will eat you alive.