Physix Frontier · News Briefing Card (enbrief · Sep 18, 2026)

Zhipu AI's GLM Self-Optimizes Inference, Boosting Throughput 3x

KEY FACTS

  • The Zhipu GLM team disclosed a recursive self-improvement practice driven by the GLM-5.3 model via an Infra Agent.
  • This agent built production-grade inference services from scratch on a cluster of over 100,000 domestic chips.
  • In less than two weeks, it increased end-to-end throughput to three times the initial baseline while matching mainstream GPU cost efficiency.
  • The anonymous model Ox-Alpha surpassed 62 trillion token calls within six days of launch and is already in real-world use.

KEY DATA

>100,000 unitsCluster Size
3xThroughput Increase
<2 weeksOptimization Time
62 trillionToken Calls (6 Days)

PHYSIX OBSERVATION

This marks a shift for large models from being 'deployed' to 'self-operating.' AI is no longer just a code generator but a closed-loop participant with systems engineering capabilities. Against the backdrop of domestic computing power, using algorithmic optimization to bridge hardware gaps and drastically lower per-token costs is a key path to breaking Nvidia's ecosystem monopoly. The engineering realization of this 'self-evolution' may reshape the paradigm of building inference infrastructure.

Source: enbrief original report ↗