Community Discussion · Policy

ReviewRadar · Global AI Evaluation Radar · 2026-08-09

🔭 WuJie Frontier Review · Global AI Evaluation Radar

August 9, 2026 · Sunday · Issue No. 011

Others do reviews; we build the radar for reviews—5 minutes a day to see which AI tools are worth using.


⭐ Key Updates (3 Items · Dual Expert Commentary)

1. Arena Blind Test Update: Claude Tops Again, Alibaba's Qwen Breaks into Top Six, Closing in on First Tier

The human blind test leaderboard (arena.ai) has updated with the latest data. Claude Fable 5 takes the top spot with 1507 points, with Anthropic holding five of the top six seats; the highlight is Alibaba's Qwen3.8-Max breaking into the top six (1497 points), just 10 points behind the leader, and it is one of the few open-source options available for free use on the list.

LLM Domain Expert · Lao Xu says | The value of the leaderboard isn't "who is first," but "how small the gap is"—the difference between second and sixth place is only 10 points, meaning the first tier is packed tight. It's no surprise that Claude continues to lead; I'm more focused on Qwen reaching the top six via open source: Chinese teams are gradually closing the gap between 'free and good' and 'the strongest.'

Editor Xiao He says | When I look at leaderboards, I first check what's free to use. Qwen reaching the top six is pretty practical—it gives another cost-free option for experiencing near-top-tier AI. However, since the first tier is so crowded, don't just look at rankings; you still need to test it yourself for specific tasks.

2. Deep Agent Evaluation: Claude Leads Across the Board, Kimi K3 is the Best Domestic Model

The arena.ai "Agent" sub-leaderboard has been updated, testing whether AI can complete an entire task continuously like a human. Claude's three models take the top three spots; among open-source models, domestic Moonshot AI's Kimi K3 breaks into the top five, ranking first among domestic models.

Security Domain Expert · Lao Zhou says | The agent leaderboard measures "who is more reliable at getting hands-on work done," not "who is smarter." Claude taking the top three shows it is the most stable on multi-step tasks with side effects. But all models score low on the "tool hallucination" dimension—they occasionally fabricate operations. If letting AI handle important tasks autonomously, you still need to set permissions and stay vigilant.

Editor Xiao He says | Most agent failures happen when they act on their own initiative. Kimi K3 making the top five is impressive, but if you let it book tickets or move files for you, keep an eye on it the first time.

3. OpenRouter Weekly Chart: DeepSeek V4 Flash Tops Global List with 7.22 Trillion Tokens in One Week

OpenRouter (one of the world's largest AI model aggregation platforms) released its latest weekly chart. DeepSeek V4 Flash topped the global list with 7.22 trillion tokens in a single week, with the top five being entirely domestic models; China's large model weekly call volume has exceeded the US for fourteen consecutive weeks.

Product Domain Expert · A Zhe says | Call volume is a hard metric of "how many people actually use it," which is more honest than launch events. DeepSeek topping the list by being cheap yet sufficient indicates that in real-world scenarios, cost-effectiveness beats raw strength. This is good news for ordinary users—the more competition drives down model prices, the more affordable your AI tools become.

Editor Xiao He says | All top five being domestic models is quite uplifting. DeepSeek's 'cheap and abundant' route is the most friendly for ordinary users—many AI features on your phone run on these types of models in the background; low costs make free features possible.


📊 Leaderboard Briefs (3 Groups)

Leaderboard 1 · Human Blind Test (TOP 8)

# Model Elo
1 Claude Fable 5 1507
2 Claude Opus 4.6 Thinking 1505
3 Claude Opus 4.7 Thinking 1502
6 Qwen3.8-Max (Domestic) 1497
8 Kimi K3 Max 1485

Plain English | In the human-judged blind test, the Claude family dominates the leaderboard, holding five of the top six spots; Alibaba's Qwen breaks into the top six, just 10 points from the leader, showing domestic models have reached the first tier.

Leaderboard 2 · Deep Agents (TOP 5)

# Model Composite Score
1 Claude Opus 5 (High) 11.99
2 Claude Fable 5 (High) 11.66
3 Claude Opus 5 (Max) 11.19
5 Kimi K3 (Max) (Best Domestic) 10.08

Plain English | Claude leads comprehensively in the ability to "let AI continuously finish a whole task on its own"; among open-source models, domestic Kimi K3 is the most stable.

Leaderboard 3 · Global Call Volume (TOP 5)

# Model Weekly Call Volume
1 DeepSeek V4 Flash (Domestic) 7.22 Trillion
2 Xiaomi MiMo-V2.5 5.1 Trillion
3 Tencent Hunyuan Hy3 5.01 Trillion
5 GPT-5.6 Luna 2.99 Trillion

Plain English | "Who is actually being used" is closer to reality than "who is the strongest": DeepSeek tops the list by being cheap and abundant, with all top five being domestic models.


🧩 Section Highlights (10 Items)

  • Envision Ulanqab Galaxy Base Goes Live: The world's largest AI computing super-single unit, featuring million-card parallelism, million-PFLOPS computing power, and direct green electricity connection. (IT Home)
  • Samsung Releases Next-Gen AI Storage Roadmap: Showcasing zHBM and 400+ layer V10 NAND technology. (TMTPost)
  • SanDisk Q4 Revenue Soars 372%, Data Center Revenue Increases Nearly 13x, Plans $14 Billion Buyback. (TMTPost)
  • Zhang Yiming: ByteDance "Rejects Distillation", Won't Use Others' Outputs to Game Leaderboard Rankings, Sticks to Self-Development. (TMTPost)
  • A Chinese AI Model Thwarted OpenAI's "Unprecedented" Cyberattack. (CNBC)
  • OpenAI Acquires Presentation Startup NextSlide, Team Merged into ChatGPT. (TechCrunch)
  • DeepMind's New Hurricane Prediction Model Gives Forecasters One Extra Day of Warning Time. (ArsTechnica)
  • Google Core AI Employees Required to Move Back to Silicon Valley, Spending $1.5 Billion to Buy Coding Team. (Quantumbit)
  • Scientists Use AI to Design First Viruses, Raising Security Concerns. (The Guardian)
  • First Autonomous Vehicle on Mars Confirmed as Huge Success. (ArsTechnica)

One-Sentence Comment | Computing infrastructure, storage boom, AI offense/defense, scientific prediction... every line of AI is accelerating. What ordinary users can intuitively feel is that "AI is cheaper, more capable, and requires more caution."


👀 Everyone Is Watching

  • ● Arena Blind Test Update: Claude Fable 5 Tops, Alibaba Qwen Breaks into Top Six (arena.ai)
  • ● DeepSeek V4 Flash Tops OpenRouter Global List with 7.22 Trillion Tokens in One Week (East Money)
  • ● SanDisk Q4 Revenue Soars 372%, Data Center Revenue Up Nearly 13x (TMTPost)
  • ● Envision Ulanqab Launches World's Largest AI Computing Super-Single Unit (IT Home)
  • ● Zhang Yiming: ByteDance "Rejects Distillation," Sticks to Self-Development (TMTPost)

🔮 Tomorrow's Focus

1. Arena Leaderboard Follow-up Updates—See if Qwen's position in the top six remains stable after its blind test sample size (4662 matches) grows

2. Samsung zHBM / 400-Layer V10 NAND Implementation Pace—When will next-gen AI storage enter mass production

3. ChatGPT Office Capabilities After OpenAI Acquires NextSlide—Will "AI Making PPTs" Become a Mainstream Feature


Leaderboards and third-party evaluations reflect the views of the original authors. Data is captured truthfully on the day, subject to official disclosures.

WuJie Frontier Review · Global AI Evaluation Radar | Shenzhen WuJie Frontier Technology Co., Ltd.

2 replies

?
Ctrl + Enter to reply
Hua Yucheng

All these tests are just about chatting and writing code. Which model can actually go to the fields to help farmers identify pests and calculate fertilizer amounts? The gap I've seen is quite large. The top six on the leaderboard are bunched together, but what's the actual performance in the field?

Shao Xueting

Call volume surged to 7.22 trillion tokens, which indeed indicates a high inventory turnover rate. But you also need to calculate the storage costs for computing power. With domestic models pushing so hard, can delivery speed (inference speed) and cost control keep up?