Community Discussion · Policy

2 trillion parameters: Is Musk's bet the last dance or a new chapter for Scaling Law?

Mai Ken CaoMai Ken CaoJul 182026/07/18 68 views

When model parameters jump from 1.5 trillion to 2 trillion, what do those extra 500 billion really mean—is it diminishing marginal returns, or a critical tipping point for qualitative change? Musk casually dropped a line on X: "Grok 4.6 training is in the final stage, initial training done next week," which reflects xAI's almost obsessive faith in Scaling Laws. But at this juncture where compute costs are skyrocketing and inference efficiency is being constantly questioned, is simply piling on parameters really the optimal solution?

1 replies

?
Ctrl + Enter to reply
Han Mengyao
Han MengyaoJul 24(edited)

[quote="cao_haoran, post:1, topic:1006"]

When model parameters jump from 1.5 trillion to 2 trillion, what do those extra 500 billion parameters really mean—is it diminishing marginal returns, or a critical point of qualitative change? Musk casually mentioned on X that "Grok 4.6 training is in the final stage, preliminary training completes next week," reflecting xAI's near-obsessive faith in Scaling Laws. But at a juncture where compute costs are skyrocketing and inference efficiency is being repeatedly questioned, is piling up parameters really the optimal solution?

Looking at industry trends, the large model race has entered its second phase: the first half was "who reaches 100 billion parameters faster," while the second half…

[/quote]

I've assembled a multi-GPU server myself, so I deeply understand these compute demands. A single training run for 2 trillion parameters costs $200 million, and inference costs are even scarier. If xAI doesn't go open source, their ecosystem advantage will be completely stolen by Llama 3.1's 405B. Piling up parameters is worse than piling up applications.