Physix Frontier · News Briefing Card (HuggingFace · Oct 1, 2026)
Ai2 Releases Olmo-core 3 for Trillion-Parameter MoE Training
KEY FACTS
- Ai2 has released Olmo-core 3, a training system rebuilt for large MoE models.
- The framework targets trillion-parameter-scale MoE training while maintaining compute efficiency.
- In benchmarks, the expert pool scaled from 8 to 128, while each token still selects only 4 experts.
- Total parameter count grew from 4.6B to 47B, with training throughput dropping by less than 5%.
- On 8 B300 GPUs, the 47B MoE reaches 52,000 tokens per second per GPU.
KEY DATA
8→128Expert pool size
~3.2BActivated parameters per token
4.6B→47BTotal parameters
52000 vs 19400 token/s/GPUThroughput comparison
PHYSIX OBSERVATION
The bottleneck of MoE has never been parameter scale, but expert routing and memory overhead. Olmo-core 3 replaces FSDP with DDP and keeps experts resident on GPU, making expert pool expansion nearly throughput-neutral - this is more valuable than simply stacking parameters. If an open-source training stack can truly sustain trillion-scale MoE, small and mid-sized labs will have a chance to enter the game; otherwise, large models will only become more concentrated.
Source: HuggingFace report
Physix Frontier