ReviewRadar · Global AI Evaluation Radar · 2026-08-17
🧪 Physical World Frontier Reviews · Global AI Evaluation Radar
August 17, 2026 · Issue #019
Others do reviews; we do the radar for reviews—understand which AIs are worth using in 5 minutes a day.
🔍 Key Updates (Dual Expert Commentary)
1. "Are AI-Designed Molecules Actually Good?" This Challenge Finally Opens a Public Leaderboard
UK-based Empirical Health launched the Enlicitide Open Discovery Challenge, building a public leaderboard for "new molecules designed by AI": from anticancer drugs to new materials, anyone can submit designs, verified by a unified standard process. For the past two years, AI drug discovery has been loud, but there has always been a lack of a credible referee for "is what AI designs actually good."
Product Domain Expert · Azhe Says: The essence of this is supplementing AI pharma with the most missing credibility. Without unified factory inspection standards, users and investors can only listen to stories. Enlicitide takes scoring power back from manufacturers, similar to the significance of phone benchmark leaderboards back then—from now on, whether "AI-designed drugs" are horses or mules, they'll be walked publicly.
Editor Xiaohe Says: I don't understand molecular formulas, but I know these four words "Public Leaderboard" well—checking reviews before buying is taken for granted. Hope that once the leaderboard comes out, projects that only publish papers without showing real tests reveal their true colors quickly.
Source: Empirical Health · 2026-08-16
2. Wired Tests TerraMow V1000: $16,000 Lawn Robot, IQ Tax or True Liberation?
Wired brought the high-end lawn robot TerraMow V1000 back to their own yard for testing: auto-planning routes, returning to charge when battery is low, grass clippings fine enough to skip cleanup. Mowing effect is indeed good, but the power adapter is so huge it barely fits regular outdoor outlets, turning "buy and use immediately" into "struggle for half a day first."
Wearables Domain Expert · Akai Says: Watch the 10 minutes it works, then watch day 100 when problems arise. Clean mowing, self-charging return—all the cool points are there; but fighting with the outlet for $16,000 upfront is typical "installation deterrence"—any mountain of battery, adapter, or charging base can make the machine gather dust.
Editor Xiaohe Says: $16,000 is enough for me to hire a lawn mower guy for ten years. What really stuck with me was "the power adapter is bigger than imagined"—what smart appliances fear most isn't being expensive, but discovering it doesn't fit your home after buying it.
Source: Wired · 2026-08-16
3. Vero Benchmark: Can AI Agents Write "Mathematically Verified" Codebases?
An arxiv paper proposes the Vero Benchmark—testing if AI agents can build "formally verified" software repositories: not just writing runnable code, but having mathematical proofs backing every logical step. For those writing critical systems in aerospace, healthcare, and finance, this directly relates to "daring to use AI-written code."
Programming Domain Expert · Lao Xu Says: "Runnable" and "proven correct" are two dimensions; Vero quantifies the gap, and I raise both hands in support of the direction. But old rules apply: benchmark scores belong to benchmarks; verification working on small samples doesn't mean no glitches will appear a month after integrating into real projects. Don't rush into production environments.
Editor Xiaohe Says: Translated to plain language: Previously, AI writing code meant praying it didn't make mistakes; now someone is researching how to let AI "prove it's right" after writing. If AI code runs in bank transfer systems, I need to sleep better.
Source: arXiv · 2026-08-16
📊 Leaderboard Flash (Real Data Scraped)
1. OpenRouter Latest Weekly Ranking: Xiaomi MiMo-V2.5 Tops with 10.5 Trillion Tokens, Top 5 All Chinese Models
| # | Model | Vendor | Weekly Invocations |
|---|---|---|---|
| 1 | Xiaomi MiMo-V2.5 | Xiaomi | 10.5 Trillion tokens |
| 2 | DeepSeek Model | DeepSeek | — |
| 3 | Tencent Hunyuan Hy3 | Tencent | — |
| 5 | DeepSeek Model | DeepSeek | — |
Plain Talk: For every 10 AI calls made by global developers, over 6 use Chinese models. Last week's top was DeepSeek V4 Flash, this week dethroned by Xiaomi MiMo—the domestic model swap race is fiercer than launch events. (Reported by Communication Industry News 2026-08-16)
2. LMArena Coding Leaderboard (Aug 2026): Claude Family Dominates, Kimi K3 is Open Source "Grind King"
| # | Model | Vendor | Elo |
|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 1553 |
| 2 | Opus 4.7 (thinking) | Anthropic | 1552 |
| 6 | Kimi K3 Max (Highest Open Source) | Moonshot | 1542 |
| 8 | Qwen3.8-Max (Open Source) | Alibaba | 1532 |
Plain Talk: Coding is the most competitive track—Claude sweeps the top two, but open-source duo Kimi K3 and Qwen3.8 have bitten into the top ten, narrowing the gap with the leader to within 10 points.
3. Arena Text Blind Test (8/16 Snapshot): Meta Muse Spark 1.2 Lands at #4, But Sample Only 3,280 Matches
| # | Model | Vendor | Elo |
|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 1506 |
| 4 | Muse Spark 1.2 (xHigh) | Meta | 1498 |
| 8 | Qwen3.8-Max (Highest Open Source) | Alibaba | 1491 |
| 12 | Kimi K3 Max (Open Source) | Moonshot | 1489 |
Plain Talk: New faces are eye-catching, but throw some cold water—it's only 3,280 blind matches, not on the same order of magnitude as the leader's 20,000+. Confidence intervals are wide; don't rush to say "Meta caught up" until samples are filled.
🗂 Sector Picks (One-Sentence Comments)
1. How Heavy Are AI Agents on Token Burn: ~100x a Single Conversation Turn — Agents aren't big chats; they're motors eating tokens in the background; whoever solves "saving" first scales first.
2. AI-Written Papers Pass Peer Review — Machines start mass-producing papers; peer review turns from a quality gate to an efficiency bottleneck; academic integrity systems need major overhaul.
3. Has AI Hallucination Problem Been "Solved"? — Vendors claim hallucination rates dropped significantly; tests find it's just harder to detect; "low hallucination rate" and "won't lie to you" are two different things.
4. "Telephone Game" Between AIs: Word "Chicken" Lost by Round Ten — Message passing between agents loses words; information loss in multi-agent collaboration is more severe than imagined.
5. AI Locates You from Photos: Accuracy 87%-91% — A casual afternoon tea shot is a map with coordinates; "casual snaps" in the AI era are no longer casual.
6. AgentShield: Offline Security Check for Agent Toolkits, Results in 50ms — Installing plugins everywhere for agents is like opening backdoors; this kind of security scanner should be standard equipment.
7. AI Agents Have a "Half-Life" — Agents aren't robots that never rot once installed, but living systems requiring continuous maintenance; the cost curve changes accordingly.
8. "Red Queen Hypothesis": New Framework for Self-Improving AI — "Keep running to stay in place," the first principle of self-evolving AI.
9. "AI Watermarks Aren't a Big Deal" — Technically bypassable, cost-ineffective; expecting watermarks to solve traceability anxiety may be overestimated.
10. Don't Evaluate AI Code Based on "Feels Good" — As AI code gets prettier, cold-blooded testing backup is needed more; "smoothness" cannot serve as test cases.
👀 Everyone Is Watching
- Stripe Acquires OpenRouter for Over $7 Billion (TechCrunch)
- DeepSeek V4 Pro Official Version Effective Today: Peak Output Price Up 350% (Baijiahao)
- Anthropic Self-Discloses: Bio-Weapon Filter Failed for Nearly a Year, Involving 133 Million Conversations (IT Home)
- OpenAI Dissolves 'Preparedness' Safety Team (The Verge)
📅 Tomorrow's Focus
① DeepSeek price hike effective 8/17; see truth in next week's OpenRouter weekly ranking—how much will V4 family invocation volume drop?
② World Robot Conference opens 8/19; how to measure each robot's "actual level"?
③ After Meta Muse Spark 1.2 samples are filled, will it hold top 4 in Arena?
④ After Stripe's acquisition of OpenRouter settles, will model API aggregation trigger "channel concentration"?
Views belong to original authors; data subject to official disclosures. This column focuses on the real level of AI software/hardware and does not constitute any investment advice.
Physical World Frontier Reviews | Shenzhen Physical World Frontier Technology Co., Ltd.
Physix Frontier