Physix Frontier · News Briefing Card (IT Home · Sep 17, 2026)
Zhipu GLM AI builds own inference system, boosting throughput 3x
KEY FACTS
- The Zhipu GLM team disclosed its recursive self-improvement practice, driven by the Infra Agent powered by GLM-5.3.
- This agent built a production-grade inference service from scratch on a cluster of over 100,000 domestic chips.
- Within less than two weeks, it increased end-to-end throughput to three times the initial baseline, achieving costs comparable to mainstream GPUs.
- The anonymous model Ox-Alpha surpassed 62 trillion token calls within six days of launch and is already in real-world use.
KEY DATA
>100,000 unitsCluster Size
3xThroughput Increase
<2 weeksOptimization Time
62 trillion6-Day Token Calls
PHYSIX OBSERVATION
This marks a shift for large models from being 'deployed' to 'self-operating.' AI is no longer just a code generator but a closed-loop participant with systems engineering capabilities. In the context of domestic computing power, using algorithmic optimization to bridge hardware gaps and drastically reduce per-token costs is a key path to breaking Nvidia's ecosystem monopoly. The engineering realization of this 'self-evolution' may reshape the paradigm for building inference infrastructure.
Source: IT Home report
Physix Frontier