20x compute gap: how should models actually be tested
Community Discussion · Tracks

20x compute gap: how should models actually be tested

Sister Liang on ValuationSister Liang on ValuationSep 112026/09/11 61 views

I spent two days running the same set of questions through different large models. Conclusion first: It’s worth doing if you’re preparing to use AI for product, content, or research, but don’t take it directly as an investment conclusion. The title says compute power differs by 20x; in my tests, what matters is how far apart cheap, stable, and deliverable results are. The ceiling of this track depends on how big the business gets; the competitive landscape depends on real usage, price, and ecosystem, not just launch events.

If you’re completely new to this, let me explain the jargon simply. Large models are AI programs that read questions and write answers. Tokens are the text units used for billing and counting—you can roughly think of it as charging by word count. Context is how much conversation it can remember at once; beyond that, it forgets. Compute power is the machine resources needed to run it; news saying compute power differs by 20x refers to underlying resource gaps, not necessarily answer quality differing by 20x. Inference is the actual process of generating answers.

The task is simple: build a minimal test sheet. Open your usual large model webpage, click model selection, choose a domestic open-source model (AI where parameters and code are partially public), then click new chat. Enter the first question: Ask the model to explain tokens in one sentence, assuming the reader knows nothing about tech. Hit send. You expect to see a Chinese explanation. If the page shows usage or API billing (where fees are deducted per call), take a glance; if not, just record whether the answer made sense to you.

For the second question, switch to a real scenario: Ask the model to judge if the headline is exaggerated. The headline is "Compute Power Differs by 20x, How Do Chinese Large Models Counterattack?" Here, follow up once, asking it to give three judgments: valid, invalid, and insufficient evidence. You expect the model to start breaking down the problem rather than chanting slogans along with the headline. For the third question, test delivery: Ask it to compress the above judgment into under 50 characters. If it’s too long, ask it to re-compress, keeping only the conclusion and one reason. Record the model, answer, number of follow-ups, and usability.

I used Insightify for a week to organize fields into a table, and later found the most useful field was how many rounds it took me to get a usable result. Cheap doesn't mean effortless; effortlessness is closer to real cost.

On day one, I made the rookie mistake of comparing different models using different questions. One asked for copywriting, another for code; the results looked good in their own ways but weren't comparable. On day two, I fixed the same set of questions and the same number of follow-ups, and only then did the differences emerge.

Another pitfall is treating free quotas as cost. Web versions often don't show token consumption; you think it's cheap, but once context lengthens, costs rise. I later tested capability with short contexts first, then stability with long contexts. You can paste a long text in segments, ask it to list key facts first, then mark potential out-of-context interpretations. Last week, when writing about route expansion, I said don't just search keywords; build detailed tables. Testing models is the same.

Also, models will try to please you. If you ask if Chinese models are counterattacking, it might say yes; if you ask if OpenAI is being overtaken, it might say yes too. I previously emphasized cross-validation with multiple models, and it came in handy again. In my tests, if the answer flips immediately when you rephrase the same question, it’s not reliable enough for direct judgment yet.

I’ve seen recent reports discussing long-term winners in China's AI large models, placing Zhipu and DeepSeek in the foundation model category; there’s also news mentioning that after the triple release of GLM 5.2, Kimi K3, and DeepSeek V4, the market reignites debates periodically.

Reports are reports; you only know the feel by testing it yourself. Price wars will change the competitive landscape. OpenAI lowering prices for lower-tier models—will the low-price advantage of domestic models be squeezed? Keep watching.

Next step: Try your own work tasks. Take product manuals, customer service dialogues, news headlines, and run them through the same table. Models with commercial value shorten processes. If next week models drop prices again or open-source more, can this test sheet still reveal differences? You’ll need to run another round.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts