Community Discussion · Tracks

Tokens on Shelves: Zhipu Is Selling Billing, Not Just Access

Bili GeBili GeSep 62026/09/06 46 views

Listing Tokens on Tmall is like moving wholesale oil tickets from a refinery to a convenience store shelf. Previously sold to enterprises were projects, deployments, and negotiations. Now displayed are Lite, Pro, Max, and Team Seats, with monthly fees of 118, 538, 1078, and 598. This retail shift turns compute costs into SKUs that can be priced, renewed, and compared.

Zhipu's flagship store GLM Coding Plan is based on GLM-5.3, compatible with over 20 mainstream agentic coding tools like ZCode, ClaudeCode, and Codex. After purchase, users get API call quotas, free to use within limits, with excess charged at standard API rates.

My first reaction focused on revenue structure. H1 2026 revenue was 954 million RMB, up 399.7% YoY. Open platform and API revenue was approx. 825 million RMB, accounting for 86.5%, while localized deployment revenue dropped to 13.5%. This shows Zhipu is less like the old model-selling/project-doing company and more like a platform selling call volume. How high the technical barrier is depends on the model, but even more so on the bill.

Selling APIs on the official site differs from selling packages on Tmall. The official site is developer self-service: pay-as-you-go, flexible, long decision cycle. Tmall is a retail shelf, translating complex Token consumption into monthly subscription packages, lowering the understanding threshold. With fewer than 2000 followers and sales not yet displayed, short-term conversion may be limited. But placing AI subscriptions alongside video memberships and office software will slowly train user habits.

The points system is worth examining closely. Input, output, and cache hits convert to points, with different coefficients for different models and MCP capabilities, plus off-peak discounts. AI Agent costs are hard to fix.

One task might split into planning, retrieval, coding, debugging, and replay, burning Tokens at every step. Whoever can turn consumption into predictable, auditable ledgers has pricing power. I wrote about Agent energy consumption a few days ago, thinking Agent products need to make uncertain energy usage into verifiable accounts. Zhipu selling packages now moves somewhat in this direction.

However, valuation shouldn't just look at growth rates. Subscription cash flow looks pretty, but gross margin is another matter. Token packages are like phone plans; heavy users easily exhaust quotas, and low-price acquisition can lead to losses from high usage. Cache hits, off-peak discounts, and overage charges are all about controlling unit costs. The question is whether prices cover inference compute, channel commissions, customer service, and refunds. If renewal rates don't go up, single-Token contribution margin breaks even or goes negative, and the 86.5% API revenue share just shifts risk from project-based to subscription-based.

Some analyses suggest AI subscriptions have high ARPU, no logistics, and renewals, making them suitable for Tmall shelves. Tmall also released guidelines for AI software and application listings in April, requiring more transparent token quotas.

For Tmall, this is a new high-value virtual category. For Zhipu, Tmall isn't the main battlefield; developers and enterprise workflows are. Selling Tokens is more like advertising. Whether ads convert to retention depends on embedding into daily actions like coding, debugging, refactoring, and auditing. Supporting 20+ tools is borrowing others' entry points. If toolchains later develop metering, budgets, permissions, and logs, the barrier rises.

Looking at exit paths, I care more about monthly active developers, package renewal rates, and single-Token contribution margin. 399.7% revenue growth is fierce, but hard tech valuation ultimately returns to unit economics. Models can catch up, prices can race to the bottom. What's hard is getting customers to include Tokens in fixed budgets, reducing ad-hoc reimbursements.

This listing pushes LLM sales from project-based to subscription cash flow, and compute from black box to auditable usage. Next, watch renewals and margins.

2 replies

?
Ctrl + Enter to reply
Old Luo
Old LuoSep 7

Production line takt times are strictly fixed. No matter how smart your model is, high inference latency drags down the entire line. Have you calculated the integration costs?

Brother Kun

Selling tokens is like selling tons of cement. Stop with the fluff and just tell me how many kilowatt-hours it takes to run a model?