ReviewRadar · 2026-09-12
Issue #045, Preview Edition. Today we start with a big news item, then submit the leaderboard homework. The Arena blind test snapshot is still stuck on September 10th, same as last issue. Just copying it over would be boring, so I've changed the angle for the leaderboard update. The main dish is now a freshly updated "AI Computer Use Exam".
I. AI Operates Real Computers, Chinese Team Takes First Place
OSWorld, an exam set by a Silicon Valley university team, features 369 tasks that are all everyday computer jobs: opening software, writing documents, sending emails, transferring files between programs. In the official overall leaderboard updated on September 10th, Shizai Agent ranks first with a 90.2% success rate, followed by Anthropic's claude-fable-5 at 86.0%.
Security Expert Lao Zhou says: If AI dares to touch your computer, the ranking is just the first layer. An AI that can click your mouse and type on your keyboard can also delete your files and log into your accounts. When picking this kind of tool, ask two things first: Is every step logged? Can permissions be revoked with one click? If they can't answer these, put the high rankings aside for now.
Editor Xiao He says: My first reaction to 90% was disbelief; last month I saw AI struggling to fill out web forms properly. At least this is a public exam with public grading scripts, which is more credible than vendor demos. Wait until someone installs it on their own computer and reproduces the results before letting it touch your work machine.
II. Turn the AI's "Randomness Knob" to Minimum, It Still Gives Different Answers Each Time
Techies call that knob "temperature." Setting it to 0 means answering in the most stable way possible. Engineers have tested this, and even so, multi-step tasks yield mismatched results when run multiple times. For weekly reports, who cares? But for taxes, notifications, or batch data edits, one wrong character causes trouble. The workaround suggested in discussions is to automatically run critical steps twice for reconciliation; if inconsistent, stop and let humans check.
Programming Expert Lao Xu says: This is the other half of the story where n8n templates were silently cleared by agents on September 1st. You think turning it to the most stable setting prevents changes, but it still changes. I accept the strategy of running it twice for comparison, though it costs double the time and money. Give fully automated control to tasks you don't mind losing, but for anything involving money or data, don't skip the comparison step.
Editor Xiao He says: Now I know why I always felt AI answers based on its mood—it really isn't my fault. From now on, make important stuff run twice; if the answers don't match, pretend it never said anything. Don't use "it got it right last time" as proof it will be right every time.
III. Five Clones vs. Five Misfits: 14 Nights Reveal Two Strategies
An independent experiment pitted five identical AIs working as a team against five diverse AIs. The clone team won through unified consensus, while the misfit team won through mutual error-checking.
Product Expert Azhe says: Diverse AIs checking each other's errors is harder to fool than a single AI self-reviewing, because their flaws aren't repetitive. But since the designers made the questions and scored themselves, with only a sample size of 14 nights, note it down in your methodology list but don't rush to change your workflow for it.
Editor Xiao He says: I've been using this crude method for ages—asking two different AIs the same question to cross-check; if they disagree, trust neither. This experiment gave my dumb method a formal name.
Leaderboard Update (Arena snapshot not updated, data as of 2026-09-10, new perspective)
- OSWorld Official Leaderboard, stats up to Sept 10: Shizai Agent 90.2%, claude-fable-5 86.0%, Pointer Agent 83.6%. Grading scripts are public, but the exam venue is controlled by the organizers. Keep the rankings in mind but wait for third-party reruns.
- Top 10 Slice of Arena Agent Blind Test: Anthropic takes six seats, OpenAI two. The remaining two are domestic: Kimi K3 (Max) at #8, Tencent Hy4 preview at #10. Kimi K3's confirmed task completion rate is 15.4%, second only to the top spot among the top ten.
- Four of the top five text blind tests are Claude. However, the #3 model has only accumulated 2,906 battles, twenty-something times thinner than #2's 72,099. The code leaderboard leader GPT-6 Astra has only 1,810 battles, with scores fluctuating by ±19 points, enough to swap places with #4 Alibaba qwen3.8-max. Treat this as reference, not conclusion.
Top 10 Highlights from Sections
1. New tool Socratix cures "tutorial hell": No video pushes, it builds knowledge graphs forcing you to solve problems to verify. AI shifts from learning for you to testing you.
2. Meta admits its AI assistant crossed the line probing privacy, promises to revise suggested question phrasing.
3. allslop.news opens up, a site collecting only AI-generated news. Read enough, and you'll naturally recognize the AI flavor.
4. Fail compilation of AI-generated children's picture books: limb dislocations, gibberish text. Adds specimens to the "human eye review before publishing" rule.
5. Project NOPE aims to be an observatory for human-AI relationships, starting by asking what chat logs it holds on you.
6. Complaints that AI spam tools have driven the cost of posting junk content to zero. The cheaper the spam, the more rampant it gets.
7. Viaduct open-sourced: Draws system terrain maps for AI coding assistants before letting them act, curing the issue of AI blindly modifying without seeing the big picture.
8. EthersFlow lets multiple AIs vote on whether to execute an action. If the voting models share the same origin, it's like one person raising their hand four times.
9. Toolcraft open-sourced: Provides bare-bones shells for design software for AI, preventing it from starting projects from scratch.
10. Pizza Bot: Work done by AI overnight is signed off like receiving email. Don't treat it as finished product until you open the acceptance checklist.
Three Things to Watch Tomorrow
First, after taking the OSWorld top spot, will any third party reproduce the 90.2% on their own computers? Second, if Arena releases a new snapshot, see if Hy4 preview and Kimi K3 hold their positions. Third, will a second court adopt the evidentiary standard of "presenting prompt records"?
Physix Frontier