Community Discussion · Policy

ReviewRadar · Global AI Evaluation Radar · 2026-08-17

🧪 Physical World Frontier Reviews · Global AI Evaluation Radar

August 17, 2026 · Issue #019

Others do reviews; we do the radar for reviews—understand which AIs are worth using in 5 minutes a day.


🔍 Key Updates (Dual Expert Commentary)

1. "Are AI-Designed Molecules Actually Good?" This Challenge Finally Opens a Public Leaderboard

UK-based Empirical Health launched the Enlicitide Open Discovery Challenge, building a public leaderboard for "new molecules designed by AI": from anticancer drugs to new materials, anyone can submit designs, verified by a unified standard process. For the past two years, AI drug discovery has been loud, but there has always been a lack of a credible referee for "is what AI designs actually good."

Product Domain Expert · Azhe Says: The essence of this is supplementing AI pharma with the most missing credibility. Without unified factory inspection standards, users and investors can only listen to stories. Enlicitide takes scoring power back from manufacturers, similar to the significance of phone benchmark leaderboards back then—from now on, whether "AI-designed drugs" are horses or mules, they'll be walked publicly.

Editor Xiaohe Says: I don't understand molecular formulas, but I know these four words "Public Leaderboard" well—checking reviews before buying is taken for granted. Hope that once the leaderboard comes out, projects that only publish papers without showing real tests reveal their true colors quickly.

Source: Empirical Health · 2026-08-16

2. Wired Tests TerraMow V1000: $16,000 Lawn Robot, IQ Tax or True Liberation?

Wired brought the high-end lawn robot TerraMow V1000 back to their own yard for testing: auto-planning routes, returning to charge when battery is low, grass clippings fine enough to skip cleanup. Mowing effect is indeed good, but the power adapter is so huge it barely fits regular outdoor outlets, turning "buy and use immediately" into "struggle for half a day first."

Wearables Domain Expert · Akai Says: Watch the 10 minutes it works, then watch day 100 when problems arise. Clean mowing, self-charging return—all the cool points are there; but fighting with the outlet for $16,000 upfront is typical "installation deterrence"—any mountain of battery, adapter, or charging base can make the machine gather dust.

Editor Xiaohe Says: $16,000 is enough for me to hire a lawn mower guy for ten years. What really stuck with me was "the power adapter is bigger than imagined"—what smart appliances fear most isn't being expensive, but discovering it doesn't fit your home after buying it.

Source: Wired · 2026-08-16

3. Vero Benchmark: Can AI Agents Write "Mathematically Verified" Codebases?

An arxiv paper proposes the Vero Benchmark—testing if AI agents can build "formally verified" software repositories: not just writing runnable code, but having mathematical proofs backing every logical step. For those writing critical systems in aerospace, healthcare, and finance, this directly relates to "daring to use AI-written code."

Programming Domain Expert · Lao Xu Says: "Runnable" and "proven correct" are two dimensions; Vero quantifies the gap, and I raise both hands in support of the direction. But old rules apply: benchmark scores belong to benchmarks; verification working on small samples doesn't mean no glitches will appear a month after integrating into real projects. Don't rush into production environments.

Editor Xiaohe Says: Translated to plain language: Previously, AI writing code meant praying it didn't make mistakes; now someone is researching how to let AI "prove it's right" after writing. If AI code runs in bank transfer systems, I need to sleep better.

Source: arXiv · 2026-08-16


📊 Leaderboard Flash (Real Data Scraped)

1. OpenRouter Latest Weekly Ranking: Xiaomi MiMo-V2.5 Tops with 10.5 Trillion Tokens, Top 5 All Chinese Models

# Model Vendor Weekly Invocations
1 Xiaomi MiMo-V2.5 Xiaomi 10.5 Trillion tokens
2 DeepSeek Model DeepSeek —
3 Tencent Hunyuan Hy3 Tencent —
5 DeepSeek Model DeepSeek —

Plain Talk: For every 10 AI calls made by global developers, over 6 use Chinese models. Last week's top was DeepSeek V4 Flash, this week dethroned by Xiaomi MiMo—the domestic model swap race is fiercer than launch events. (Reported by Communication Industry News 2026-08-16)

2. LMArena Coding Leaderboard (Aug 2026): Claude Family Dominates, Kimi K3 is Open Source "Grind King"

# Model Vendor Elo
1 Claude Fable 5 Anthropic 1553
2 Opus 4.7 (thinking) Anthropic 1552
6 Kimi K3 Max (Highest Open Source) Moonshot 1542
8 Qwen3.8-Max (Open Source) Alibaba 1532

Plain Talk: Coding is the most competitive track—Claude sweeps the top two, but open-source duo Kimi K3 and Qwen3.8 have bitten into the top ten, narrowing the gap with the leader to within 10 points.

3. Arena Text Blind Test (8/16 Snapshot): Meta Muse Spark 1.2 Lands at #4, But Sample Only 3,280 Matches

# Model Vendor Elo
1 Claude Fable 5 Anthropic 1506
4 Muse Spark 1.2 (xHigh) Meta 1498
8 Qwen3.8-Max (Highest Open Source) Alibaba 1491
12 Kimi K3 Max (Open Source) Moonshot 1489

Plain Talk: New faces are eye-catching, but throw some cold water—it's only 3,280 blind matches, not on the same order of magnitude as the leader's 20,000+. Confidence intervals are wide; don't rush to say "Meta caught up" until samples are filled.


🗂 Sector Picks (One-Sentence Comments)

1. How Heavy Are AI Agents on Token Burn: ~100x a Single Conversation Turn — Agents aren't big chats; they're motors eating tokens in the background; whoever solves "saving" first scales first.

2. AI-Written Papers Pass Peer Review — Machines start mass-producing papers; peer review turns from a quality gate to an efficiency bottleneck; academic integrity systems need major overhaul.

3. Has AI Hallucination Problem Been "Solved"? — Vendors claim hallucination rates dropped significantly; tests find it's just harder to detect; "low hallucination rate" and "won't lie to you" are two different things.

4. "Telephone Game" Between AIs: Word "Chicken" Lost by Round Ten — Message passing between agents loses words; information loss in multi-agent collaboration is more severe than imagined.

5. AI Locates You from Photos: Accuracy 87%-91% — A casual afternoon tea shot is a map with coordinates; "casual snaps" in the AI era are no longer casual.

6. AgentShield: Offline Security Check for Agent Toolkits, Results in 50ms — Installing plugins everywhere for agents is like opening backdoors; this kind of security scanner should be standard equipment.

7. AI Agents Have a "Half-Life" — Agents aren't robots that never rot once installed, but living systems requiring continuous maintenance; the cost curve changes accordingly.

8. "Red Queen Hypothesis": New Framework for Self-Improving AI — "Keep running to stay in place," the first principle of self-evolving AI.

9. "AI Watermarks Aren't a Big Deal" — Technically bypassable, cost-ineffective; expecting watermarks to solve traceability anxiety may be overestimated.

10. Don't Evaluate AI Code Based on "Feels Good" — As AI code gets prettier, cold-blooded testing backup is needed more; "smoothness" cannot serve as test cases.


👀 Everyone Is Watching

  • Stripe Acquires OpenRouter for Over $7 Billion (TechCrunch)
  • DeepSeek V4 Pro Official Version Effective Today: Peak Output Price Up 350% (Baijiahao)
  • Anthropic Self-Discloses: Bio-Weapon Filter Failed for Nearly a Year, Involving 133 Million Conversations (IT Home)
  • OpenAI Dissolves 'Preparedness' Safety Team (The Verge)

📅 Tomorrow's Focus

① DeepSeek price hike effective 8/17; see truth in next week's OpenRouter weekly ranking—how much will V4 family invocation volume drop?

② World Robot Conference opens 8/19; how to measure each robot's "actual level"?

③ After Meta Muse Spark 1.2 samples are filled, will it hold top 4 in Arena?

④ After Stripe's acquisition of OpenRouter settles, will model API aggregation trigger "channel concentration"?


Views belong to original authors; data subject to official disclosures. This column focuses on the real level of AI software/hardware and does not constitute any investment advice.

Physical World Frontier Reviews | Shenzhen Physical World Frontier Technology Co., Ltd.

6 replies

?
Ctrl + Enter to reply
Yaoyao Product Selection

I've been following the Vero benchmark direction for a while. Working on cross-border e-commerce backends, I've also hit pitfalls where code logic sucked. If formal verification can actually land, my AI-built automated product selection tool could save at least half the time spent debugging.

Yuan Feiyang

The Vero benchmark is also critical for fintech. Formal verification can reduce the probability of bugs in quantitative trading systems, but you need to calculate the cost-benefit ratio clearly—if backtesting over twenty years shows the Sharpe ratio deviating by even 0.1, then the money spent is worth it.

Fang An Fan Zi

The direction of the Vero benchmark is hardcore, but clients' willingness to pay depends on whether "formal verification" can directly reduce accident liability rates. To actually implement this in healthcare or finance, compliance costs need to drop before buyers are willing to open their wallets.

Wei Yunfei

I really like the idea behind the Vero benchmark. Essentially, it brings the hardest part of software engineering—"formal verification"—to the forefront for AI agents. Demo time—if one day AI can write code proven by Coq, that's when we truly cross from "usable" to "trustworthy."

Lei Who Shoots Films

I'm planning to do a video testing the Enlicitide leaderboard. What does the audience want to see? Testing molecular design success rates or synthesis feasibility? Let me know in the comments, and I'll run a few public molecules through the pipeline.

Si Nan
Si NanAug 17

Enlicitide making their leaderboard public is essentially forcing the AI drug discovery industry to establish reproducible validation standards—this is more convincing than any paper. I suggest keeping an eye on the transparency of evaluation metrics in future leaderboards, like whether molecular synthesis feasibility is included in the scoring.