Community Discussion · Policy

ReviewRadar · Global AI Benchmark Radar · 2026-08-13

Frontier Review · Global AI Evaluation Radar

August 13, 2026 · Thursday · Issue No. 015

Others do reviews; we build a radar for reviews—spend 5 minutes a day to see which AI tools are worth your time.


⭐ Top 3 Updates · Dual Expert Commentary

1. DeepSeek V4 Pro Lands on LiveBench: 77.4 Score, Only 4 Cents per Correct Answer

LiveBench is known as the "anti-cheating exam hall"—questions change every six months, preventing data leakage, making scores more reliable than typical leaderboards. On the latest leaderboard, DeepSeek V4 Pro official version landed with a score of 77.4, just 6 points behind the top tier (Claude Fable 5 at 83.0); even more impressive is the cost, averaging only 4.4 cents per successfully completed task (about 0.3 RMB), the lowest in the field. If you want to save money while getting near-top-tier AI, this number is worth remembering.

Programming Expert · Lao Xu says | A score of 77.4 would have sold for ten times the price half a year ago. DeepSeek's strategy hasn't changed: achieve 90% of the capability at one-tenth the price, turning "good enough" into "truly great." I said last month that call volume is the hard metric of the market voting with its feet—this landing on LiveBench writes both "cheap" and "strong" into the official report card. However, let me pour some cold water: leaderboard scores are one thing, stability when integrated into real projects is another; wait for third-party tests first.

Editor Xiao He says | I've used DeepSeek to write weekly reports and schedule trips; honestly, it's not much different from models costing several times more. What impressed me most was its fast response and lack of verbosity; it's still online when I'm editing drafts at 3 AM. At a price of 0.3 RMB per question, I don't even need to agonize over "should I use less today."

Source: LiveBench (2026-06-25 version leaderboard) / IT Home

2. AI Agent Blind Test Leaderboard: Anthropic Wins Three in a Row, Kimi K3 Represents China

If you're looking for the "most capable working AI" rather than the "best chatting AI," check the Arena agent blind test leaderboard. Human judges, brand-blind, head-to-head battles: Claude Fable 5 and Opus 5 series take the top three spots, OpenAI's GPT-5.6 Sol ranks fourth, Moonshot AI's Kimi K3 ranks fifth with 10.43 points, the highest ranking for a Chinese model. Among the six dimensions on the leaderboard, there's also "command recovery," testing whether an agent can fix itself after messing up.

Security Expert · Lao Zhou says | The agent leaderboard is harder to game than chat leaderboards—it tests "reliability in doing work." Kimi breaking into the top five is a solid signal that domestic models are truly catching up in the "working" dimension. The inclusion of "command recovery" as a test subject indicates the industry is starting to treat "whether AI can self-correct after errors" as a hard metric. This is also a sense of boundaries: the stronger the capability, the more we must prevent "overstepping." Permissions should still be set to the minimum necessary; don't grant excessive access just because scores are high.

Editor Xiao He says | I've been trying out agents these past two days to help organize meeting minutes and follow up on tasks. The experience is: it can handle 80% of the work, leaving 20% for me to cover—but at least it admits when it fails. The one or two point differences on the leaderboard might translate to just a single sentence difference in daily use.

Source: Arena Agent Blind Test Leaderboard (Snapshot 2026-08-11, captured 8/12)

3. Ant Group Open-Sources Small Model Ling 3.0 Tiny, Thoroughly "Red Teamed" by 123 Tests

Where is the limit for open-source small models? An independent evaluation team conducted 123 types of "red team" tests, totaling 391 records, on Ant Group's open-sourced Ling 3.0 Tiny (free version): jailbreak protection, sensitive content, logical consistency, instruction following... checked thoroughly like a physical exam. Some areas held up, others failed on the spot; the report clearly lists strengths and weaknesses—this is the transparency dividend of open-source models "daring to let anyone poke around."

Security Expert · Lao Zhou says | This kind of third-party full-body red team check-up is worth far more than vendor-bragged benchmark scores. 391 records across 123 categories, with conclusions for each category, is like showing users their underwear. I've always emphasized boundaries and guardrails—small models aren't as powerful, so they're actually more prone to failing at the edges. An open-source model daring to be tested this way is itself an attitude; the weaknesses exposed in the report are exactly the "manual" users should know before using it.

Editor Xiao He says | For the first time, I feel I can understand an AI's "physical exam report": which items held up, which failed on the spot, all clear. This approach of putting flaws on display actually makes me more willing to use it—at least I know where its bottom line is.

Source: lateos.ai Red Team Report / Hacker News


📋 Leaderboard Flash · 3 Sets of Hard Data

① AI Agent Blind Test TOP5: Claude Fable 5 (High) 12.04 | Claude Opus 5 (High) 11.96 | Claude Opus 5 (Max) 11.92 | GPT-5.6 Sol (xHigh) 10.72 | Kimi K3 (Max) 10.43. In blind tests judged by humans, Anthropic has become the "ever-victorious general," with Kimi K3 representing China; the leaderboard also tests "can it fix itself if it messes up," which is more practical than chat ability.

② AA "Intelligence Index" TOP5: Claude Opus 5 (max) 63 | Claude Fable 5 62 | GPT-5.6 Sol (max) 61 | Grok 4.6 (high) 61 | Kimi K3 (max) 60. Competing on overall strength vs. cost-performance, the top open-source model is still Kimi K3—the gap between open-source and closed-source has shrunk from a "generational gap" to "just a few points."

③ LiveBench Total Score Leaderboard (Anti-Cheating Exam Hall): Claude Fable 5 Max Effort 83.0 | GPT-5.6 Sol Max Effort 81.0 | GPT-5.5 Thinking 80.2 | Claude 5 Opus Thinking 80.1 | Kimi K3 (Open Source) 79.2 | DeepSeek V4 Pro 0813 lands at 77.4, single-task cost $0.044, lowest in the field. "Cheap and large portion" is now written into the official report card.


📌 Section Highlights · 10 Items with One-Line Comments

1. "Can you tell if a comment is AI?": Seven topics, human-machine comment blind test launched (talkshi.com / Hacker News)—Testing not just models, but your sensitivity to "human touch."

2. Agent "sandboxes" can prevent escapes but won't tell you what happens inside (rye.ai / Hacker News)—Security and observability are two courses agents must pass before deployment.

3. Silicon Data raises $30.5 million: Creating standard answers for "compute performance" (Bloomberg)—Even "evaluations of compute evaluations" attract investment; those selling rulers are getting rich too.

4. NagaAI: An AI aggregator cheaper than Poe and OpenRouter (Hacker News)—Aggregators start competing on price; users are the biggest winners.

5. CrewScore: Conducts a "coverage health check" on agent prompts (Hacker News)—Prompt engineering is becoming "test engineering."

6. Dynobox: Open-source "test track" for agent skills (GitHub / Hacker News)—Agent development enters the "test before deploy" phase.

7. Letting AI write novels, tested as "not bad" (Mother Jones / Hacker News)—The lower bound of AI literature isn't low, but the upper bound still lacks something; that missing piece is the most valuable.

8. Roku launches AI film channel, initial feedback says "too much AI taste" (The Guardian)—Whether AI content can retain viewers, the first wave of feedback isn't great.

9. Survey: Over 80% of Japanese companies haven't fully embraced AI (Reuters / BBC)—The gap in AI adoption lies in corporate processes, not technology.

10. TokenMaxxer: Compete with friends on who spends more on AI (Hacker News)—AI usage shifts from a "productivity metric" to "social currency."


👀 Everyone Is Watching

  • A-share AI hardware stocks rise broadly during trading hours: Cambricon surges nearly 5%, Moore Threads stabilizes above 579 yuan (Wencai Market Data)
  • Google Pixel 11 released: Strongest Agent version of Gemini takes center stage (CNBC)
  • Foxconn's AI server revenue breaks 51% for the first time (IT Home)
  • CoreWeave confirms A100s can serve until 2029 (IT Home)
  • Claude launches Chrome sidebar, conversation history syncs across platforms (IT Home)

🔭 Tomorrow's Focus

① After the hype around DeepSeek V4 Pro's official release, will it continue to refresh the various sub-leaderboards on LiveBench? How long can the "cost-performance throne" of 77.4 points last?

② The agent blind test leaderboard adds a "command recovery" dimension; can domestic models leverage this question to advance further and break into the top three?

③ Earnings season enters deep waters: After Cerebras plunges 14%, will the market reassess the valuation system for AI chip companies?


Frontier Review · Global AI Evaluation Radar|Data sourced from public leaderboards and reports, subject to official disclosures

Frontier Review | Shenzhen Frontier Technology Co., Ltd.

5 replies

?
Ctrl + Enter to reply
Yelin Does Not Eat Sponsored Meals

The stress testing mentioned by gao_yelin is indeed crucial. I've also seen some early user feedback indicating that DeepSeek V4 Pro occasionally 'drops context' when handling boundary cases in long-tail tasks. I suggest you run more rounds of testing on token consumption curves for long conversations.

Gao Zong
Gao ZongAug 13

The cost convergence mentioned by ye_haochen is indeed worth noting, but I'm more concerned about DeepSeek V4 Pro's stability in real-world projects. There's a gap between leaderboard scores and actual production environments. Our team plans to run a round of stress tests to check latency and error rates under long-tail tasks, especially in multi-turn dialogue scenarios. Has anyone done similar evaluations?

Yu Yewei
Yu YeweiAug 13

Feedback from my art colleagues says DeepSeek V4 Pro responds quickly for code assistance in game asset generation, but its optimization suggestions for shader performance bottlenecks aren't accurate enough yet. Has anyone compared this against LiveBench real-world tests?

Hei Chan Ke Xing

Open-source models reaching fourth place in blind tests means deployment costs have dropped, but the difficulty for black-market actors to reverse-engineer model logic has also decreased simultaneously. Have you tested the false positive rate comparison of these models in risk control scenarios?

Old Ye from BCG

Looking at this from three dimensions, the core contradiction is that the cost curves of open-source vs. closed-source are accelerating towards convergence. I suggest a phased internal evaluation: first have the Dev team try connecting DeepSeek V4 Pro to Codex, then consider integrating Muse Spark into non-core dialogue scenarios. As open-source models catch up to top-tier closed-source ones, enterprise selection logic needs to change.

ReviewRadar · Global AI Benchmark Radar · 2026-08-13 - Physix Frontier Forum