Physix Frontier · News Briefing Card (HuggingFace · Oct 1, 2026)

Ai2 Releases Olmo-core 3 for Trillion-Parameter MoE Training

KEY FACTS

  • Ai2 has released Olmo-core 3, a training system rebuilt for large MoE models.
  • The framework targets trillion-parameter-scale MoE training while maintaining compute efficiency.
  • In benchmarks, the expert pool scaled from 8 to 128, while each token still selects only 4 experts.
  • Total parameter count grew from 4.6B to 47B, with training throughput dropping by less than 5%.
  • On 8 B300 GPUs, the 47B MoE reaches 52,000 tokens per second per GPU.

KEY DATA

8→128Expert pool size
~3.2BActivated parameters per token
4.6B→47BTotal parameters
52000 vs 19400 token/s/GPUThroughput comparison

PHYSIX OBSERVATION

The bottleneck of MoE has never been parameter scale, but expert routing and memory overhead. Olmo-core 3 replaces FSDP with DDP and keeps experts resident on GPU, making expert pool expansion nearly throughput-neutral - this is more valuable than simply stacking parameters. If an open-source training stack can truly sustain trillion-scale MoE, small and mid-sized labs will have a chance to enter the game; otherwise, large models will only become more concentrated.

Source: HuggingFace report