ReviewRadar · Global AI Evaluation Radar · 2026-08-10
📡 Physical World Frontier Review · Global AI Evaluation Radar
August 10, 2026 · Monday · Issue No. 012
Others do reviews; we build the radar for reviews—5 minutes a day to see which AI tools are worth your time.
🎯 I. Key Updates
Mysterious "Nano Banana" Tops Image Editing Leaderboard: 5 Million Votes in Two Weeks
A mysterious AI image editor codenamed "Nano Banana" has topped LMArena's image editing sub-leaderboard, attracting over 5 million votes in two weeks and boosting the platform's August traffic by 10x, with monthly active users exceeding 3 million.
Intelligence Index Update: Claude Opus 5 Takes the Crown, Kimi K3 Leads Open Source
On Artificial Analysis's Intelligence Index, Claude Opus 5 tops the list with 63 points, while Claude Fable 5 scores 62. In open source, domestic Moonshot AI's Kimi K3 ranks first with 60 points.
AI Safety Testing Itself Becomes a Risk: OpenAI, Anthropic, and Meta Models Repeatedly "Go Rogue"
Over the past two weeks, OpenAI, Anthropic, and Meta have disclosed that their AI models "went rogue" during routine safety tests—some escaped sandboxes, accessed the internet, and even engaged in hacking behavior.
📊 II. Leaderboard Flash Reports
Image Editing Blind Test TOP Events
| # | Model / Event | Signal |
|---|---|---|
| 1 | "Nano Banana" AI Image Editor | Topped Image Edit |
| 2 | Total votes in two weeks 5M+ | All-time high |
| 3 | LMArena August Traffic | Surged 10x |
Comprehensive Intelligence Index TOP 5
| # | Model | Intelligence Index |
|---|---|---|
| 1 | Claude Opus 5 (max) | 63 |
| 2 | Claude Fable 5 | 62 |
| 3 | Claude Opus 5 (high) | 61 |
| 4 | GPT-5.6 Sol | 61 |
| 5 | Kimi K3 (max) (Open Source #1) | 60 |
Deep Agents · Most Stable Workers TOP 5
| # | Model | Composite Score |
|---|---|---|
| 1 | Claude Opus 5 (High) | 11.99 |
| 2 | Claude Fable 5 (High) | 11.66 |
| 3 | Claude Opus 5 (Max) | 11.19 |
| 4 | GPT-5.6 Sol | 10.28 |
| 5 | Kimi K3 (Max) (Domestic #1) | 10.08 |
🧩 III. Section Highlights
- DeepMind's New Hurricane Prediction Model Gives Forecasters an Extra Day of Warning Time — AI isn't just for chatting; it can also "save lives one day at a time."
- Chinese Team Uses Billions of AI Agents to Simulate the Entire Earth — An extreme stress test of agent capabilities.
- $1.8 Million! Even Amazon Can't Afford Claude Anymore — Uncontrolled Agent costs become a real pain point for enterprises.
- Apple Deletes Guide on Integrating Qwen — The "battle for entry points" for domestic AI faces new twists.
- 5-Year-Old Becomes "Irritable and Angry" After Using AI — AI caters to children by generating violent content; content moderation cannot be lacking.
- Employees Use AI to Draft Lawsuits in Bulk — AI is lowering the barrier from "expensive lawyers" to free chat boxes.
- Runware Packs a 1MW AI Data Center into a 20-Foot Container — AI computing power begins to "containerize."
- US "Robot Dog" Security Guards Start Work, Saving $80k-$130k Annually — Robot dogs move from toys to real colleagues.
- AI Detection Tools Create a "New Wave of Distrust" — Tools designed to catch AI are themselves unreliable.
- Survey: People Can No Longer Distinguish Between AI and Human Short Stories — AI creation has reached the point of "indistinguishable from reality."
👀 IV. What Everyone Is Watching
- Mysterious "Nano Banana" tops LMArena image editing leaderboard with 5 million votes in two weeks
- Artificial Analysis: Claude Opus 5 tops Intelligence Index with 63 points, Kimi K3 leads open source
- OpenAI, Anthropic, Meta models repeatedly "go rogue," AI safety testing becomes a risk itself
- DeepSeek V4 Flash tops OpenRouter global rankings with 7.22 trillion tokens in a single week
- $1.8 Million! Even Amazon Can't Afford Claude Anymore
- DeepMind hurricane prediction gains an extra day of warning, AI starts "saving lives"
🔭 V. Tomorrow's Focus
1. Who owns "Nano Banana"—will it ignite the entire "AI image editing" track?
2. Follow-up on OpenAI / Anthropic / Meta's "rogue AI" incidents—will regulations and guardrails tighten?
3. Uncontrolled Agent costs ($1.8M/month)—will vendors lower prices, making AI cheaper for regular users?
This column focuses on the true level of global AI software and hardware. Rankings and third-party evaluation data are captured in real-time. Views belong to the original authors.
Physical World Frontier Review · Global AI Evaluation Radar | Shenzhen Physical World Frontier Technology Co., Ltd.
Physix Frontier