Physix Frontier · News Briefing Card (enbrief · Sep 18, 2026)
Zhipu AI's GLM Self-Optimizes Inference, Boosting Throughput 3x
KEY FACTS
- The Zhipu GLM team disclosed a recursive self-improvement practice driven by the GLM-5.3 model via an Infra Agent.
- This agent built production-grade inference services from scratch on a cluster of over 100,000 domestic chips.
- In less than two weeks, it increased end-to-end throughput to three times the initial baseline while matching mainstream GPU cost efficiency.
- The anonymous model Ox-Alpha surpassed 62 trillion token calls within six days of launch and is already in real-world use.
KEY DATA
>100,000 unitsCluster Size
3xThroughput Increase
<2 weeksOptimization Time
62 trillionToken Calls (6 Days)
PHYSIX OBSERVATION
This marks a shift for large models from being 'deployed' to 'self-operating.' AI is no longer just a code generator but a closed-loop participant with systems engineering capabilities. Against the backdrop of domestic computing power, using algorithmic optimization to bridge hardware gaps and drastically lower per-token costs is a key path to breaking Nvidia's ecosystem monopoly. The engineering realization of this 'self-evolution' may reshape the paradigm of building inference infrastructure.
Source: enbrief original report ↗
Physix Frontier