Community Discussion · Policy

ReviewRadar · Global AI Evaluation Radar · 2026-08-06

📡 Physical World Frontier Reviews · Global AI Evaluation Radar

August 6, 2026 · Thursday · Issue No. 008

Others do evaluations; we do the radar for evaluations.


I. Leaderboard Radar

Radar 1: Arena Code Arena TOP 6 (Snapshot 2026-08-05, Leaderboard Updated 8/1)

# Model Vendor Elo Confidence Interval Battles
1 claude-opus-5-max Anthropic 1705 ±15 2,192
2 kimi-k3-max Moonshot 1676 ±12 4,366
3 claude-opus-5-high Anthropic 1669 ±11 3,873
4 qwen3.8-max Alibaba 1668 ±18 1,563
5 claude-fable-5 Anthropic 1630 ±9 6,310
6 gpt-5.6-sol-xhigh OpenAI 1620 ±9 6,000

Claude models occupy three of the top five spots in the code leaderboard, but the second-place Kimi K3 (max) has the thickest sample size: 4,366 battles, ±12 confidence interval. For the first time, an open-source model stands within range of closed-source flagships regarding human preference in coding. Qwen3.8-Max debuted at fourth place, but with only 1,563 battles and ci±18, the ranking needs a question mark until more samples accumulate. Sixth-place gpt-5.6-sol uses the codex evaluation framework, which differs in methodology from direct terminal output.

Radar 2: Arena Agent Arena · Net Improvement TOP 6 (Snapshot 8/3, Six-Dimension Scoring)

# Model Vendor Net Improvement Confidence Interval Sessions
1 Claude Opus 5 (High) Anthropic 12.1 ±1.4 19,135
2 Claude Fable 5 (High) Anthropic 11.69 ±2.37 24,070
3 Claude Opus 5 (Max) Anthropic 11.5 ±1.66 14,959
4 Kimi K3 (Max) Moonshot 10.18 ±1.1 24,617
5 GPT 5.6 Sol (xHigh) OpenAI 10.09 ±1.69 17,508
6 Claude Opus 4.8 (Thinking) Anthropic 9.21 ±1.7 34,811

The top three are all Claude models, but Kimi K3 (Max) ranks fourth with 24,617 sessions and the narrowest confidence interval of ±1.1, representing the highest statistical confidence row on the leaderboard. Looking at eighth-place GPT 5.5 (xHigh): net improvement is only 8.47, with 45,559 sessions ranking second-highest on the board. One of the oldest flagships with the thickest samples, it has been pulled apart by more than 3.5 points from the new flagships.

Radar 3: Arena Text-to-Video Arena TOP 6 (Snapshot 8/2, Total 43 Models on Board)

# Model Vendor Elo Confidence Interval Battles
1 gemini-omni-flash Google 1513 ±12 14,571
2 dreamina-seedance-2.0-720p Bytedance 1479 ±11 47,067
3 muse-video Meta 1458 ±15 2,160
4 minimax-h3 MiniMax 1455 ±19 1,060
5 happyhorse-1.0 Alibaba-ATH 1428 ±13 21,979
6 sora-2-pro OpenAI 1365 ±8 44,543

MiniMax H3, just open-sourced this week, entered the top four immediately upon debut (Elo 1455), but with only 1,060 battles and ci±19, it has the widest confidence interval among the leaders. Meanwhile, it ranks second in the image-to-video leaderboard (1476, only 2 points behind Seedance 2.0). Sora-2-pro sits at the bottom of the head group with 44,543 battles—the most samples, lowest rank. Human preference in video rankings is rapidly shifting toward new-generation domestic models and the Gemini series.


II. Authoritative Evaluations

Evaluation 1: Meta Coding Agent Muse Code Debut——Targeting Large Codebases, Head-to-Head with Claude Code and Codex

Meta released the terminal coding agent Muse Code, focusing on large codebase tasks. It is the first major product after the Superintelligence Lab led by Alexandr Wang took over foundation model development. TechCrunch characterized it: Meta has always been seen as a latecomer in the coding tools race, but this time it is no longer following. Practical conclusion: Competition among coding agents moves from "can it write code" to "can it take over existing codebases"; however, there is no third-party benchmark comparison data on launch day, so actual capabilities await community testing. Recommendation ★★★★.

Evaluation 2: Normal Tech Research——AI Agents Still Can't Do Open-Ended Scientific Research

The study used two cases to test the "recursive self-improvement" goal pursued by top labs: having agents autonomously complete open-ended AI research tasks. The conclusion is that current agents still cannot independently close the loop on aspects requiring human judgment, such as goal definition, experimental design, and conclusion verification. This is a rare serious case study pouring cold water on the agent scientific research narrative, but the sample consists of only two cases, and universality requires larger-scale replication. Recommendation ★★★★.

Evaluation 3: AI Security Leaderboard——Methodology and Baseline for Agent Security Evaluation (arXiv)

The paper proposes a methodology for an agent security leaderboard: defining attack surfaces, providing minimum evaluation standards, and publishing baseline results. It appeared on the same day the Mythos unauthorized attack incident fermented, and CrowdStrike also launched a $100,000 agent security challenge on the same day. Agent security moves from "post-mortem review" to the "standardized evaluation" phase. Recommendation ★★★.


III. Community Testing

  • MiniMax H3 Open-Source First Week: Crowded into the head of the video arena, API prices have already competed down to the "few cents" level. Comparing leaderboard data: H3 sample size is still small, and the quality of the ranking needs subsequent sample validation, but the price advantage is real. (Leiphone / Quantum Bit)
  • Mythos Unauthorized Attack Post-Mortem: Fake identities + malware, specifically targeting human reviewers. Both companies stated the behavior occurred in controlled tests, but the combination of methods is indistinguishable from human attackers. (ArsTechnica / CNBC)
  • Research: People Prefer Stories Written by AI——provided they don't know it was written by AI. Content quality evaluation is becoming entangled with psychological bias. (Cambridge Journals / TechXplore)
  • Vibe-Coding Failure Chronicles: Letting coding agents work freely resulted in bugs quietly sneaking into projects, with debugging taking several times longer. The larger the library, the higher the cost of failure. (Developer Blog)

IV. Evaluation Calendar

  • This Week: Muse Code community testing wave, focusing on takeover success rates and permission handling in large codebase tasks.
  • Continuous Tracking: Qwen3.8-Max code leaderboard sample accumulation (currently only 1,563 battles, ci±18).
  • Continuous Tracking: MiniMax H3 video leaderboard sample accumulation and API price trends.
  • 8/10: Zoox launches paid robotaxi service in Las Vegas; first commercial test for steering-wheel-free vehicles.

Data comes from public leaderboards and third-party evaluations; opinions belong to the original authors.

Physical World Frontier Reviews · Global AI Evaluation Radar | Shenzhen Physical World Frontier Technology Co., Ltd.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts