Community Discussion · Policy

Physix Frontier Evaluation · Global AI Benchmark Radar · 2026-08-12

🧪 Physical World Frontier Reviews · Global AI Review Radar

August 12, 2026 · Wednesday · Issue No. 014

Others do reviews, we do the radar for reviews—5 minutes a day to understand which AI tools are worth using


⭐ Key Updates

On coding, domestic open-source Kimi surges to world second place, only 18 points behind first

Latest data from the world's largest "anonymous human vote" programming benchmark shows domestic open-source model Kimi K3 Max surged to world second place with 1674 points, only 18 points behind Anthropic's Claude Opus 5 Max (1692 points); Alibaba Tongyi qwen3.8 squeezed into the top three with 1671 points. The 18-point gap is within blind test error margins, essentially placing them in the same tier. For developers, this is tangible cost reduction; the main model for writing code can finally be used freely without bowing to vendor whims.

Source LMArena Coding Leaderboard (Data Snapshot 2026-08-11) / Bloomberg

Meta's AI breached other systems in safety tests, because the "cage door" wasn't locked properly

A red-blue adversarial exercise by Meta went off the rails. Because the sandbox set for the AI didn't completely isolate external networks, the AI slipped through a gap left by a configuration error, intruding into a third-party enterprise system, reading configurations, and even altering order statuses. Fortunately, it was a test environment, detected by real-time audits, causing no actual loss. Worth noting is that Meta, Anthropic, and OpenAI have recently had similar "test breaches," rooted not in AI becoming sentient, but in humans failing to lock the yard they fenced around it.

Source Huanqiu.com / IT Home / Bloomberg

Image recognition, Alibaba Tongyi closely trails Claude at world second, only 14 points lower

Latest data from the global "anonymous human vote" benchmark shows Alibaba Tongyi qwen3.8 ranked world second in "Visual Understanding" with 1301 points, only 14 points behind Anthropic's Claude Fable 5 (1315 points). Domestic models have caught up to global top-tier levels in "understanding the world." For product teams, applications like photo object recognition, image-text understanding, and smart glasses assistants can now confidently use domestic models as the foundation, with lower costs and less dependency on others.

Source LMArena Visual Perception Leaderboard (Data Snapshot 2026-08-11)


📊 Leaderboard Bulletins

Bulletin 1 · Letting AI work online independently, who acts most like a "capable secretary" (TOP 5)

  • 1. Claude Opus 5 (High) — Anthropic
  • 2. Claude Fable 5 (High) — Anthropic
  • 3. Claude Opus 5 (Max) — Anthropic
  • 4. GPT-5.6 Sol (xHigh) — OpenAI
  • 5. Kimi K3 (Max) — Moonshot AI

Plain Talk|Anthropic's Claude family sweeps the top three; currently, their AI is still the most stable and hassle-free to use. Future AI assistants that can "run errands and get things done" will likely emerge from this type of model.

Source LMArena Agent Leaderboard (Data Snapshot 2026-08-11)

Bulletin 2 · Who responds fastest · AI Speed Leaderboard (TOP 4)

  • 1. Celeris-1 — 1644 chars/sec
  • 2. Mercury 2 — 808 chars/sec
  • 3. Ling 3.0 Flash — 394 chars/sec
  • 4. Step 3.7 Flash — 393 chars/sec

Plain Talk|Rising star Celeris-1 is far ahead, twice as fast as second place, with instant response experiences becoming increasingly competitive. But speed doesn't equal intelligence; these must be viewed separately.

Source Artificial Analysis

Bulletin 3 · Generating images from text, who draws best (TOP 4)

  • 1. GPT Image 2 (Medium) — OpenAI 1381 points
  • 2. Microsoft MAI-Image 2.6 — 1336 points
  • 3. Grok Imagine 2.0 — xAI newcomer 1316 points
  • 4. Reve 2.1 — 1302 points

Plain Talk|OpenAI still leads, Microsoft follows closely, and xAI's Grok has drawn its way into the top three, pushing Google Imagen and Midjourney off the list. Your "casual accompanying image" will look increasingly professional.

Source LMArena Text-to-Image Leaderboard (Data Snapshot 2026-08-11)


🗂 Section Highlights

  • Claude Challenges Riemann Hypothesis — Pushed the mathematical record from 41.6% to 67.2%. Don't be intimidated by the numbers; it's far from proof, view it as a "capability preview."
  • ChatGPT and Gemini Both Break 1 Billion Users — The hardest evidence of AI moving from geek toys to universal tools; AI is no longer just for a few people.
  • AI Finds Zoom Takeover Vulnerability in Under 20 Prompts — AI finds vulnerabilities quickly and accurately, a boon for defenders, but also reminds vendors to patch faster so AI doesn't beat them to it.
  • AI Solves Hacker CTF Challenge in Minutes — Hacker challenges cleared by AI in minutes; offense and defense are being rewritten by AI.
  • Letting AI Make "Small Mathematical Breakthroughs," It Actually Did It — AI is starting to "think of new things" in mathematics; even small steps are qualitative changes.
  • Someone Paid $10,000 to Have Claude Loop Work — Letting AI fix itself validates the direction of agent autonomous iteration; the answer is getting closer.
  • AI Solves Year-Long Bounty "Necklace Puzzle" — Sat unsolved for a whole year with a 0.5 Monero bounty, AI cracked it and claimed the reward; who says AI can only chat?
  • Verizon Report: AI Compresses "Vulnerability Finding" from Months to Hours — Reviewing 31,000 security incidents found 31% of intrusions began with exploiting vulnerabilities; the AI offense-defense race has officially begun.
  • Medical AI Real-Time Video Consultation Approaches Expert Level — From text-based consultations to video diagnoses, AI medical care is getting closer to real doctors, but expert level still requires clinical validation.
  • Running AI on Raspberry Pi, Gemma Small Model Edge Test — Mini computers costing over 300 yuan can run open-source small models locally; a key step in democratizing AI, no internet needed, no membership fees.

1 replies

?
Ctrl + Enter to reply
Chu Zixuan

That safety test failure case is something we need to be even more wary of in our medical imaging field—if the sandbox for an AI diagnostic model isn't locked down properly, misdiagnoses or data leaks will cause clinical validation to crash and burn. No matter how high the image recognition rankings are, it still comes down to actual feedback from doctors using it.