ReviewRadar · 2026-08-25 · Preview of Issue 27
Physix Frontier Review · 2026-08-25 · Issue 027 Preview
Highlights This Issue Google's mid-tier model hits three major AI blind test leaderboards overnight. An anonymous model that refuses to reveal its identity burned through 4 trillion tokens in one day, ending DeepSeek's 56-day reign on the usage leaderboard.
Google's Mid-Tier Rookie Gemini 3.7 Flash Hits Three Leaderboards Simultaneously
Arena is a blind testing ground where thousands of users anonymously vote to score AI. Strength isn't determined by vendors. On August 24, Google's Gemini 3.7 Flash hit three leaderboards in one go.
On the Agent Leaderboard, which tests "can AI do work for you," it ranked #20, jumping 15 places from the previous generation, standing in the same tier as Claude Sonnet 4.6 and GPT-5.6 Terra. On the web development leaderboard, it surged from #19 to #8. It also entered the top nine on the chat blind test leaderboard, scoring 1490. The brightest single metric is "getting things done in one shot," showing stability in long tasks without dropping the ball halfway.
Old Xu says the Agent Leaderboard tests long-range work capability. Assign a project, and the AI must break down steps, call tools, and recover from errors on its own. Google's mid-tier model's work stability has genuinely improved. However, leaderboard scores are just scores. When integrating into projects, try small tasks for a few days first.
Little He says previously, AI coding was all about the most expensive models. Now that a mid-tier model hits three leaderboards, cheap and good AI is getting closer. Note the name. Once it opens up, let it help build a simple webpage to test the waters.
Anonymous Model Burns Through 4 Trillion Tokens in One Day
Tokens are AI's counting units. Every question you ask and answer it gives is measured by them. On August 24, the anonymous model Ox Alpha quietly went online without official identity announcement. The Chinese circle calls it "Niu Lai" (Ox Come). Within one day, its call volume surged to #1 globally, setting a new single-day historical record.
Just the top five entry points fed it over 4 trillion tokens in one day. DeepSeek's 56-day winning streak in programming tools was ended like this. It can hold approximately 1.048 million tokens of content at once, equivalent to reading hundreds of thousands of words in one breath while remembering the beginning. It can view images and videos, and the preview period is free. Who owns it? How long will the free window last? Nothing is officially announced, but developers have voted with their feet.
Ah Zhe says it's another anonymous takeover. Last time "Nano Banana" anonymously dominated image editing leaderboards, playing out similarly. Entry points and free access matter more than names. For users, the harder anonymous models compete, the more free and good tools there are.
Little He says this time I won't guess who it is. I'll watch two things. 4 trillion tokens is real usage data. How long the free window lasts is what matters to us regular folks. Old rule: wait for third-party real-world tests first.
Leaderboard Quick Report
Agent Leaderboard (Arena Official Announcement 8/24)
- Gemini 3.7 Flash Overall Rank #20, jumped 15 places
- "Get Done in One Shot" Single Metric Rank #5, Signal +10%
- Same range as Claude Sonnet 4.6 and GPT-5.6 Terra
Web Development + Chat Leaderboard (Arena Official Announcement 8/24)
- WebDev surged from #19 to #8, Data Analysis category #7
- Text Arena Rank #9, Score 1490
Usage Leaderboard (NBD, KuaiKeJi Reports 8/24)
- Ox Alpha topped the list on Day 1, setting single-day record
- Cumulative >4 trillion tokens from top five entry points
- DeepSeek's 56-day consecutive record ended
Sector Highlights: 10 Items
- Agent Leaderboard adds hard metric "Completion Rate." Can it get done is more valuable than can it do it. Google rookie ranks #5 on this single metric.
- Google's mid-tier model stands in the same range as Claude Sonnet 4.6 and GPT-5.6 Terra. Price weight can be reduced slightly when choosing tools.
- In the web development arena, Gemini 3.7 Flash surged from #19 to #8. Data, gaming, and design categories all entered top ten.
- Chat blind test leaderboard score 1490, rank #9. For daily chatting and copywriting, mid-tier models are already top-notch.
- Ox Alpha spec sheet: holds hundreds of thousands of words at once, reads images/videos, free during preview.
- DeepSeek's 56-day winning streak ended by an unknown entity. Developers voting with feet is the most real thing.
- Identity mystery. Tokenizer fingerprint resembles Zhipu, researchers hint at Google. Who is Niu Lai? Waiting for official announcement.
- DeepSWE preliminary score ~63 (real-world exam for AI fixing code). Overall close to Claude Opus 4.8. Official stance. Waiting for third-party re-tests.
- Even Stripe's boss came to watch, saying it's very impressive. Big shot endorsements should be taken with a grain of salt.
- 1.048 million token ultra-long memory. Basically fits the entire "Three Body Problem" trilogy. Feed a whole book in and it remembers the beginning.
Everyone Is Watching
- Humanoid robot runs 100m in 9.58 seconds, tying Bolt's record (The Verge / IT Home)
- AI open-source platform Hugging Face rumored in $13 billion acquisition talks (TechCrunch)
- Xiaomi releases three self-developed chips. Xuanjie O3 first supports LPDDR6 (IT Home)
- BYD Yangwang U7 holds driver conference today, previews unprecedented extreme challenges (IT Home)
Tomorrow's Focus
1. Yangwang U7 Driver Conference debuts today. Results of mass-produced car extreme challenges?
2. Enflame Technology subscription opens Sept 2. Domestic GPU IPO window approaching.
3. Nvidia earnings imminent. Analysts say it could be the next catalyst for AI stocks.
Physix Frontier