ReviewRadar · 2026-08-23 · Issue 25 (Preview)
Physical World Frontier Review · 2026-08-23 · Issue No. 025 (Preview Version)
Others do reviews; we build radar for reviews. Spend 5 minutes daily to understand which AI tools are worth using.
Structure of this issue: Key Updates (3 items, dual-expert commentary) | Section Highlights (10 items) | Trending | Tomorrow's Watchlist
Key Updates
1. AI Civil Servant Exam: Claude Opus 5 Works as a Clerk in the Ministry of Foreign Affairs for Over a Week
What it can do for you: Let AI handle repetitive clerical tasks, reply to emails, and organize materials. Someone let Claude Opus 5 act as a "clerk" in the Ministry of Foreign Affairs for over a week, handling real official document workflows, leaving records and retrospectives. It serves as a field trial for "Can AI work within the system?" It can do plenty, but questions remain open: which systems can it touch, who is responsible for errors, and where are the boundaries?
Security Expert Lao Zhou says|"Can AI do it" and "Should AI do it" are two different things. Before deploying AI, answer the first question: permission boundaries. Which systems can it access? What decisions can it make for you? Who backs up errors? Is there full-process logging? Trials can be bold, but if letting AI touch critical information, stick to the old rule: set permissions to minimum, authorize item-by-item, don't hand over all keys at once.
Editor Xiao He says|First reaction was "Am I going to lose my job?" Then I saw it was a pilot with human oversight and auditable records. This reinforces my belief: before letting AI touch important things, clarify which doors it can open. I've walked this path; don't rush to give away all the keys.
Source: Hacker News (2026-08-23)
2. Stronger GLM-5.3 Launched, Why Didn't It Go Viral?
What it can do for you: Daily upgrades for domestic models—translation, coding, material organization—each generation stronger. Zhipu's new GLM-5.3 primarily improved coding, long-task agents, and cybersecurity capabilities by scaling training. API officially opened on Aug 19, pricing same as previous gen, stronger capability, no price hike. Retrospective shows this launch's buzz was far lower than before. Too many models, shrinking gaps; the era of "launch equals viral" is over.
Product Expert A-Zhe says|Lack of virality despite capability upgrades indicates category maturity. Users no longer get excited by "parameters got bigger," asking only if it's directly usable in their hands. GLM-5.3 is an upgrade for developers; no new C-end entry points, hence low presence. My judgment stands: Entry is Model. Whoever connects capability to places users visit daily is the one who goes viral.
Editor Xiao He says|To ordinary users, model strength depends on usability in their frequent apps. GLM has always been my "good Chinese helper." 5.3 is stronger without price hikes; noted. Virality doesn't matter; smooth usage does. Like phone OS updates—who shouts "I updated!" every day?
Source: TMTPost (2026-08-22)
3. Before AI Writes Code, Make It Submit a "Homework Sheet"
What it can do for you: Before AI writes code, clearly state "which files to modify, how to modify, and why." If mismatched, don't start. Generating code is increasingly easy; confirming consistency between code and user needs is hardest. Doing more leads to bigger mistakes. This developer added a pre-step for AI agents: analyze, write plan, get human approval, then act. The larger AI output, the more valuable this "acceptance gate."
Programming Expert Lao Xu says|Hits the root. After AI boosts code volume, verification becomes the new bottleneck, similar to the recent "7-hour test queue" article. Making AI submit homework before acting essentially front-loads verification to before pen touches paper. I agree with the direction. Note: Beautiful plans don't guarantee correct changes. After the homework sheet, humans must still monitor changes and run tests. Don't treat this step as a disclaimer.
Editor Xiao He says|Isn't this just "Show me what you're going to do before you do it"? Same as hiring renovation workers: discuss quotes and construction plans first, don't just smash walls. Watching AI work is similar: the easier it seems, the more confirmation needed beforehand, to avoid discovering wrong walls torn down after completion.
Source: Hacker News (2026-08-22)
Leaderboard Flash · Who Was Strongest This Week
Coding Blind Test Leaderboard (Snapshot as of 2026-08-22)
| Rank | Model | Vendor | Score |
|---|---|---|---|
| 1 | Claude Opus 5 (Max) | Anthropic | 1691 |
| 2 | Kimi K3 (Max) | Moonshot AI | 1674 |
| 3 | Qwen3.8-Max | Alibaba | 1669 |
| 4 | Claude Opus 5 (High) | Anthropic | 1663 |
| 5 | Grok 4.6 (High) | xAI | 1629 |
Plain English interpretation: In a coding blind test where humans judge and models quiz each other, Anthropic took first place, but Chinese vendors have closed in to second and third, less than 30 points behind the top two. Leaderboard not yet updated; data as of 2026-08-22.
AI Agent Leaderboard (Snapshot as of 2026-08-22)
| Rank | Model | Vendor | Net Improvement Score |
|---|---|---|---|
| 1 | Claude Opus 5 (High) | Anthropic | 12.5 |
| 2 | Claude Opus 5 (Max) | Anthropic | 12.0 |
| 3 | Claude Fable 5 (High) | Anthropic | 11.6 |
| 4 | Kimi K3 (Max) | Moonshot AI | 10.4 |
| 5 | GPT 5.6 Sol (xHigh) | OpenAI | 9.7 |
Plain English interpretation: This leaderboard tests whether AI can complete tasks autonomously and recover from failures. Anthropic swept the top three. Kimi K3 is the only Chinese contender in the top four. Leaderboard not yet updated; data as of 2026-08-22.
AA Intelligence Index (Official Release 2026-08-22)
| Rank | Model | Vendor | Highlight |
|---|---|---|---|
| #4 | Kimi K3 | Moonshot AI | Highest ranking currently for a Chinese vendor |
| Following closely | DeepSeek V4 Flash 0731 | DeepSeek | Newly released last week, AA score approx. 50 |
Plain English interpretation: Third-party institutions calculate an "overall IQ" for mainstream models. Kimi K3 reached global #4, the best current result for a Chinese vendor. DeepSeek V4 Flash 0731, released last week, follows closely and is the brightest spot on the "intelligence per dollar spent" chart.
Section Highlights
1. Prompts Lie: How to Create a "Fingerprint" for AI Models?
Each vendor's model has unique answering habits and traces, like human handwriting. The problem is prompts can deceive; wrapping a shell can mimic another model. The article summarizes methods to "verify identity" of models: examining output statistical features, calculating token distributions, running behavioral fingerprints. These are especially useful in anonymous blind tests and third-party gateway evaluations.
One-line comment|Verifying model identity is the foundation for third-party evaluation and transparency.
2. Scientists Set "Exam Syllabus" for AI, Six Principles to Evaluate Cognitive Ability
What exactly should AI capability evaluations test? An academic article proposes six basic principles: tasks must reflect real usage, evaluations must be robust, don't look only at total scores, prevent models from "gaming the test." It establishes a coordinate system for the chaotic AI evaluation landscape.
One-line comment|Finally, someone is seriously building a coordinate system for evaluation methodology. Good news.
3. Moving DeepSeek Coding Agent to Cloudflare for Always-On Work
Someone ported the entire DeepSeek Harness coding agent experience to Cloudflare edge computing. Personal coding agents no longer need to guard a server; open a browser and use it.
One-line comment|Coding agents moved from local to edge; the barrier to ready availability drops another notch.
4. TechSkills: Equipping AI Coding Agents with a Bag of "Skill Packs"
Open-source project TechSkills prepares modular "skill packs" for AI coding agents. General models are smart but lack specific trade crafts. Packaging these crafts for agents is like hiring a master craftsman for a novice.
One-line comment|General models equipped with domain skill packs: right path, depending on maintenance sustainability.
5. GitX: "Cleaning Up" Code Messes Made by AI
AI agents love changing code, but Git commit history becomes a mess. GitX packages Git workflows into AI skills, enabling agents to organize messy changes into clean, revertible commit records.
One-line comment|Writing code isn't enough; you must know how to clean up the scene. This skill hits daily pain points.
6. Why Do AI Products Fail? A Three-Year-Old Retrospective Still Holds Today
A retrospective started in 2023 remains fresh. The author summarizes common causes of AI product failure: unclear problem definition, lack of genuine user need, treating technology as the end goal. AI capabilities are stronger today, but reasons for failure remain the same.
One-line comment|Pitfalls in AI products don't change; product methodology is the real skill.
7. CyberStrike: Open-Source "AI vs AI" Security Attack Testing Tool
Open-source project CyberStrike is an AI toolkit for offense-defense drills, using AI to simulate attacks and repeatedly test weaknesses in own systems. Popular saying in security circles: If you don't use AI to attack your own system, opponents will use AI to attack you.
One-line comment|Turning "AI vs AI" from slogan to runnable tool. Self-testing is better than getting hit.
8. Turning Apple Website Style "Scroll Videos" into AI Skills
Someone packaged their entire workflow for creating Apple-style scroll video webpages into reusable AI agent skills. Install in agent, say a word, generate similar webpage.
One-line comment|Doing it once isn't hard; solidifying experience into skills is the real asset.
9. Knoku: AI Answers Your Questions, Every Answer Comes with Sources
Knowledge Q&A tool Knoku focuses on "answers with citations," finding answers in your own docs, files, and team knowledge bases, attaching sources for easy verification. Specifically treats AI's serious nonsense.
One-line comment|Whether answers can be traced is the first hurdle for ordinary users trusting AI.
Physix Frontier