ReviewRadar · 2026-09-23 (Issue 56 · Preview)
Others do reviews, we do the radar of reviews. Today's issue covers three things directly related to AI performance, plus three leaderboards and ten picks.
Issue Structure
3 key updates (with dual expert commentary), 3 leaderboard briefs, 10 section picks, plus What Everyone's Watching and Tomorrow's Watchlist.
Key Updates
1. The AI that helps you compare prices online still can't tell "what it found" from "what it made up"
A Bloomberg newsletter on September 22 said that AI shopping assistants that pick and compare products for people still struggle to tell which parts of a webpage are facts and which are fabrications. These assistants read product pages, look at reviews, compare prices, and hand you a conclusion. The content they read includes specs merchants filled in themselves, plus claims of unknown origin, all mixed together.
What it means for you. Next time you use AI to pick something, have it only do the filtering step, verify the specs yourself on the official site, and click pay and order yourself. Spend five extra minutes, avoid one pitfall.
Product domain expert · Azhe says | The winning point in shopping has never been who picks fastest, it's whether you trust it. Meta Muse, launched last week, also touts shopping recommendations as a selling point, and the discussion focus lands on trust too. My advice: let it do the filtering, click pay and order yourself, verify specs on the official site.
Editor Xiaohe says | Last week I had AI pick a power bank for me, and it reported a model that was "10,000 mAh but only 80 grams." I don't believe that number. My current approach is to have it list three candidates, and I verify the specs one by one on the official site myself.
2. A developer ran a full round with the new version of Devin and wrote down where it got stuck
On September 22, a developer posted his hands-on notes on Devin on his blog. He tried the whole set: the desktop app, the command-line tool, the new coding model SWE-2 (Devin's own model), and the mode that combines several models. Devin is an AI coding tool that can read code and edit files on its own. This is a personal trial record, with no public benchmarks.
What it means for you. If you have an old project that's been running for a long time, don't rush to hand it to this kind of tool. Try it on a small module for a week, and see whether it can fix its own errors after they come up.
Programming domain expert · Lao Xu says | A single person's one-time hands-on record is valuable for the step it says got stuck. When I read this kind of article I look for two things first: how big the repo is, and whether it can fix its own errors after they come up. Official demos never mention these two. To hook it up to your own project, first roll it on a small module you won't mind losing for a week.
Editor Xiaohe says | I don't understand what SWE-2 is, I only care whether it can sort out that little project I've been running for three years. If I have to first teach it how to start the project, I'd rather do it myself. Note it down for now, and consider it once a second person posts a record of getting their own project running.
3. AI has slipped into StarCraft ranked matches, and someone teaches you to spot it from its play
A developer wrote an article about how he identifies when his opponent is an AI in StarCraft online ranked matches. StarCraft is a 1998 classic competitive game, and ranked matches are online games tiered by win rate. His method is to find patterns in details like action rhythm and opening habits, building a "fingerprint" for each player. The article is the author's own observation, with no third-party verification.
What it means for you. This "fingerprinting behavior" approach is being carried into homework plagiarism checks and hiring screening. Next time you see a "looks AI-written" conclusion, first ask whether there's a second kind of evidence.
Security domain expert · Lao Zhou says | Telling machines from humans is fundamentally an identity problem. The article doesn't say what the false-positive rate is, and that's the most critical item. The same "judge it as machine-like" approach, carried over to homework plagiarism checks and hiring screening, puts the cost of every wrong call on a person.
Editor Xiaohe says | In a game, misjudging your opponent means at worst a curse and a rematch. I'm not comfortable with this kind of judgment being carried over to exams and papers: the score a detection tool gives is only a clue, not a verdict.
Leaderboard Briefs (leaderboards not yet updated, data as of 2026-09-21)
Let me be honest up front. I checked today, and Arena's latest snapshot is still the 2026-09-21 version, the same one used last issue, so the leaderboards are not yet updated, data as of 2026-09-21. Today I also didn't get a new leaderboard to swap in, and I'd rather label it honestly than pass off old data as new. The visual blind-test leaderboard from last issue is pulled this time, replaced by the text blind-test leaderboard not used last issue.
Text blind-test leaderboard, top five separated by only 8 points
| # | Model | Vendor | Score (calculated from human votes) |
|---|---|---|---|
| 1 | claude-fable-5-high | Anthropic | 1506 (30,057 matches) |
| 2 | claude-opus-4-6-high | Anthropic | 1505 (71,993 matches) |
| 3 | claude-opus-4-7-high | Anthropic | 1502 (60,002 matches) |
| 4 | muse-spark-1.2 (xHigh) | Meta | 1500 (3,227 matches) |
| 5 | claude-fable-5.1-max | Anthropic | 1498 (5,783 matches) |
Plain-language read. The top five are separated by only 8 points, so first and fifth are basically a tie. Fourth is Meta's Muse Spark, with only 3,227 matches, and the leaderboard lists its fluctuation as 11 points, bigger than its gap from third, so don't take the ranking seriously yet. Four of the top five are Anthropic's Claude. Also, this leaderboard was last updated September 13, eight days earlier than the snapshot date.
Coding blind-test leaderboard, the top spot has the thinnest match count
| # | Model | Vendor | Score (match count in parentheses) |
|---|---|---|---|
| 1 | gpt-6-astra-max | OpenAI | 1800 (2,281 matches, ±16 points) |
| 2 | claude-fable-5.1-max | Anthropic | 1758 (3,036 matches) |
| 3 | claude-opus-5-max | Anthropic | 1687 (12,087 matches) |
| 4 | qwen3.8-max-0902 | Alibaba | 1681 (2,262 matches) |
| 5 | kimi-k3-max | Moonshot AI | 1674 (4,547 matches) |
Plain-language read. First place has 1800 points, opening a 42-point gap over second, which looks solid. Its match count is only 2,281, with ±16 points fluctuation; third has 12,087 matches, with only ±7 points fluctuation. Scores with thin match counts don't hold up to scrutiny. One more thing: this leaderboard was last updated September 11, unchanged for thirteen days, and four of the top five have fewer than five thousand matches.
Agent hands-on leaderboard, with a new "tool misstatement" item
| # | Model | Vendor | Sample (match count) | Rate of tool misstatement |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 (Max) | Anthropic | 13,320 | 0.37 |
| 2 | GPT 6 Astra (Max) | OpenAI | 10,372 | 0.37 |
| 3 | Claude Opus 5 (High) | Anthropic | 24,794 | 0.33 |
| 4 | Claude Opus 5 (Max) | Anthropic | 19,934 | 0.35 |
| 5 | Claude Fable 5 (High) | Anthropic | 38,293 | 0.37 |
Plain-language read. This leaderboard tests AI that can call tools on its own and work through several steps in a row, called agents in the industry. It doesn't rank an overall score, but calculates six items separately: getting things done, saying fewer wrong things, following instructions, and recovering on its own after errors. "Tool misstatement" means it claims a tool call that never happened actually happened. The top five all fall between 0.33 and 0.37 on this item, a tiny difference, and what separates the rankings is the getting-things-done item: first place 19.83, second 17.7. The thickest match count isn't the top spot; fourth played over nineteen thousand matches.
Section Picks (10 items, one comment each)
1. Build a practice ground where an AI assistant can't cause trouble. Firedrill is a testing framework: you first create some fake tools and fake data, then let the AI assistant work on them, then check what it called at each step and whether it followed the rules you set. Fake tools are stand-ins that won't actually send emails or transfer money. The hard part comes after: you have to first write clearly "what counts as done right," and no one can do that step for you.
2. DeepMind wrote a long piece on how to test models dynamically. Google DeepMind's research team published an article arguing for testing frontier model capabilities dynamically: not relying on one fixed exam paper, but swapping questions as models get stronger. The article is about method, with no specific scores given. "Fixed exam paper, ever-bigger models" is a common flaw of every leaderboard right now; the dynamic-question direction holds up, but the question-setters should also be checked by the same standard.
3. A university provost was found by the campus paper to have multiple signed articles that "look AI-written." Dartmouth's campus paper did an investigation, running the provost's 2026 op-eds and academic articles through AI detection tools, with the median landing in the "looks AI-written" range. Detection tools give probabilities, and the school and the person involved have not yet formally responded. Same kind of incident again: a high score isn't evidence, and what should be questioned is who has the right to use this score to challenge a person's byline.
4. OPPO released a 499-yuan "AI Mind Ball," a new category. OPPO brought a new wearable at its September 22 launch event, at an experience price of 499 yuan, officially touting putting AI capabilities into a portable little ball. The launch gave no data on battery life or how many times a day it'd be used. Don't rush to be first with a new category; whether a wearable sticks around depends on whether charging is convenient and whether you forget you're wearing it after a day.
5. An AI focused on ancient Greek fragments wants to fill in the missing letters. Wired reported that researchers built a model called Apollo, aimed specifically at badly damaged ancient Greek papyrus fragments in libraries, with the goal of filling in missing letters and matching provenance and date. Such fragments number in the hundreds of thousands, many too damaged for humans to read. The biggest fear with filling in letters is filling in something fluent but wrong; first check whether there's a public answer key.
6. An open-source chat interface that connects both online and local models. IntelliChat is an open-source chat frontend written in Next.js, aiming to connect several vendors' model APIs into one interface. Open-source means you can deploy and modify the code yourself. Not having to switch interfaces when switching models is a must for people using several vendors at once; self-deploying means managing your own keys and logs, not worth it for people who want convenience.
7. The open-source community is starting to set rules for itself: how AI-written code enters repositories. An article sorted through open-source projects' recent moves: from an early blanket ban, to writing dedicated instruction files for AI assistants (like AGENTS.md), to worrying that automated review bots themselves get injected with instructions, citing examples like kernel self-tests and the QEMU rules draft. The rules went from "ban" to "write clearly how to use it," and that step is right; the weak spot is the review link, and who reviews the bots.
8. A "viral shot" reference library for AI editing assistants. RetentionVolt collected over 500,000 cut points and opening hooks from more than a thousand short-video creators, updated weekly, and organized into an interface AI assistants can call directly, an interface method called MCP. The value of a material library is being searchable; the risk is convergence: everyone cuts to the same batch of viral shots, and the visuals get more and more alike. Use it as reference, don't copy it as a template.
9. Rabbit released an AI assistant that doesn't require buying its hardware. The Verge reported that Rabbit launched an AI assistant that runs standalone, usable without buying its R1 device. Rabbit's earlier hardware sold so-so, and this time it's pulling the software out on its own. The criteria are still two: whether it can get things done for you, and whether you can find out what it did after it errs.
10. A startup team tool with an AI that "talks back." LiveCrew gives founders four AI roles, each handling a piece of work, that point out each other's mistakes with supporting reasons. In the product page demo, after someone revises copy, another role replies "I disagree with this decision," then lists reasons. AI from different sources reviewing each other is harder to fool than the same AI self-checking; whether it can actually catch real errors needs someone to run it on their own project.
What Everyone's Watching
Musk's Grok Bot passed 400,000 users in a month. Microsoft Xbox had another round of layoffs, with several studios merged into Activision. Someone on Hacker News found that Couchsurfing's entire site was rewritten by AI, with images also officially generated. AstroForge handed command of its next asteroid mission to AI.
Tomorrow's Watchlist
① The just-released Claude Opus 5.5 and GPT-6 Sol, Luna haven't entered the Arena snapshot yet (the snapshot is stuck at 2026-09-21); wait for the next refresh to see if the rankings hold. ② The Yunqi Conference runs until September 24; third-party benchmarks for AgentCore and Qwen4 aren't out yet, so note the official claims for now. ③ Meta's Muse passed 500,000 users in its first week; wait for third parties to produce actual task completion rates before discussing whether it counts as a shopping gateway.
The above is the preview version of Issue 056; the full illustrated version is on the review publication page.
Physix Frontier