ReviewRadar · Global AI Evaluation Radar · 2026-08-11
Physix Frontier Reviews · Global AI Review Radar
Tuesday, August 11, 2026 · Issue No. 013
Others do reviews; we do the radar for reviews—understand which AI tools are worth using in just 5 minutes a day.
Key Updates
The Most Used AI Globally Is Still This One: DeepSeek V4 Flash Tops Charts with 8.83 Trillion Tokens in a Single Week
OpenRouter's weekly rankings show that last week (8/3-8/9), total global large AI model call volume was 69 trillion tokens. Chinese models hit 34.25 trillion, surpassing the US for the 15th consecutive week. DeepSeek V4 Flash took the global #1 spot with 8.83 trillion; the top four were all Chinese models. The AI most used by developers worldwide right now is made by Chinese teams.
Claude Starts Embedding "Invisible Watermarks" in AI Text: Spot Machine-Written Content at a Glance Soon
Anthropic announced that new Claude models embed watermarks invisible to the naked eye within generated text, marking the content from start to finish. In the future, you'll be able to use tools to determine if a piece of text is AI-generated, helping identify AI content online and prevent deepfake-style fraud.
Five Major Universities Jointly Release Rankings: First Robot "World Model" Evaluation Results Out
Five universities jointly released the first evaluation of robot three-view world models, providing a unified scoring standard for "how robots see the world," with rankings continuously updated. Finally, there's a third-party reference for evaluating which robot is smarter, rather than just manufacturer self-praise.
Leaderboard Briefs
| Leaderboard | Top Spot | Highlights |
|---|---|---|
| OpenRouter Weekly Call Volume | DeepSeek V4 Flash 8.83 Trillion | Top 4 all Chinese models; China surpasses US for 15 consecutive weeks |
| LMArena Text-to-Image | GPT Image 2 (1340 Elo) | Microsoft MAI-Image-2.6 lands suddenly at #2 |
| AA Intelligence Index | Claude Opus 5 (63 points) | Domestic open-source Kimi K3 scores 60, becoming "Open Source #1" |
Plain Language Interpretation: Anthropic still leads in overall capability, but free and user-friendly domestic models are gradually narrowing the gap. For ordinary people: If you want to experience top-tier AI, there's another no-cost option.
Section Picks (10 Items)
- Microsoft MAI-Image-2.6 Lands Suddenly at #2 on Text-to-Image Leaderboard — Updated every three weeks, closely chasing OpenAI's GPT Image. AI art generation is shifting from a toy to a must-contest entry point for giants.
- Large AI Model Weekly Ranking: Alibaba qwen3.8-max Breaks into Top 6 — Multiple domestic models collectively gain ground; "chasing" is turning into "running alongside."
- Century-Old Riemann Hypothesis Broken by Record-High Performance from a Certain New Claude Model — Mathematical reasoning steps up again, but the model isn't public yet; treat it as a capability preview for now.
- 3B Small Model Om AI Edge VLX Makes a Comeback — Achieves precise perception of the physical world with small parameters; AI running on phones is getting smarter.
- NVIDIA Open-Sources Magpie TTS — Multilingual low-latency voice model; voice avatars are accelerating into applications.
- More Than Half of AI-Generated Patches Are Flawed — Research shows AI writes code fast but not necessarily correctly; human oversight remains indispensable.
- Research: AI Reads Text Better Than It Listens to Voice — "Can read" and "can listen" are two different things; voice AI still has a long way to go.
- New Anthropic Research: AI Hasn't Actually Gotten Stronger — Even those building AI are questioning whether evaluations truly measure progress.
- Traces of Google Gemini 3.7 Flash Exposed — Less than a month after 3.6 release, another update is coming; AI iteration speed is terrifyingly fast.
- Zhipu ZCode Fully Upgraded — Domestic coding AI also competes on "multi-agent collaboration," with competitive pricing.
What Everyone Is Watching
- DeepSeek V4 Flash tops OpenRouter with 8.83 trillion tokens in a single week; China surpasses US for 15 consecutive weeks (East Money)
- LMArena Text-to-Image Leaderboard: Microsoft MAI-Image-2.6 lands suddenly at #2 (IT Home)
- First Robot "World Model" Evaluation Released by Five Major Universities (QbitAI)
- Century-Old Riemann Hypothesis Broken by Record-High Performance from a Certain New Claude Model (QbitAI)
- Alibaba qwen3.8-max Breaks into Top 6 Overall Rankings (IT Home)
- New Claude Models Embed Invisible Watermarks Throughout, Making AI Text Identifiable (QbitAI)
- NVIDIA Open-Sources Magpie TTS Multilingual Low-Latency Voice Model (HuggingFace)
- 3B Small Model Om AI Edge VLX Achieves Precise Physical World Perception with Small Parameters (QbitAI)
Physix Frontier