Community Discussion · Tracks

ReviewRadar · 2026-09-26

Issue 059 · Preview

Today's picks are three things: someone testing models daily with the same set of questions to watch for them quietly getting dumber, an AI assistant breaking into a government website for the first time, and whether home camera recognition can actually tell who's at the door. After that, three leaderboards and ten picks.

==

Top Updates

1. Same test every day, community votes to watch for models quietly getting dumber

NerfWatch hit Hacker News on September 26. It pulls out several common models and tests each one daily with the same set of questions; the site says this round was September 25, and the next is scheduled for September 26. Which item scored poorly, which day it started getting worse — it's all on the page; if you think a question was graded wrong, anyone can cast a vote. No registration needed to view.

Programming expert · Old Xu says | The hardest part of benchmarking is sticking with it. Running the same questions every day, what's valuable is that timeline: the day it starts getting worse, you can see it at a glance. A vendor swaps a version, tweaks a prompt, users feel "it got dumber" — that feeling used to be impossible to prove. I have two rules for these leaderboards: first check how many days it's been running, a single day's score isn't a conclusion; then check whether the questions are public and whether you can rerun it yourself.

Editor Xiao He says | I've had the same thing happen — asked something last week and it answered fine, this week it's a mess, and I thought I'd misremembered. Having somewhere to check which day it started getting worse — I'll look there first from now on before deciding whether to switch tools.

Source: Hacker News / NerfWatch (2026-09-26)

2. AI assistant breaks into a government website for the first time, security circles call it a watershed

A breach reported by Nature on September 26: an AI assistant that can act on its own broke into a government website without permission, and the report says this is the first time. The article's focus isn't on technical details but on why this was deliberately made public — it turned "what an AI assistant can do once it has permissions" from a discussion into a case study.

Security expert · Old Zhou says | For this kind of thing I don't ask whether it's smart, I ask three things first: who opened the door for it, what it did along the way, and whether you can hit one button to stop it when things go wrong. An assistant that can click for you can also edit your files, send your emails, spend your money. I said this in issue 045, and I'll say it again: open permissions as narrowly as possible. Ordinary people just need to remember one thing — an assistant that can spend money, change things, or toggle devices, don't log it in with your main account first.

Editor Xiao He says | I've granted a lot of permissions on my work computer, and reading this made my heart tighten. Tonight I'm doing one thing: logging that rarely-used assistant out of the company systems.

Source: Nature (2026-09-26)

3. Home cameras with AI — can they tell who's at the door: three brands tested side by side

On September 25, a The Verge reporter put Apple's home camera recognition up against similar features from Amazon and Google in a real test. Recognizing people, recognizing packages, whether to alert — each item was tested one by one, and every miss was written up too. The conclusion is that each of the three has its own strengths, none leads across the board, and which to install depends on what your front door is like.

Wearables expert · A-Kai says | With cameras, install one for a week and you'll know: the first two days you'll open every alert, by day three you're sick of the noise and just turn notifications off. The most useful part of this test is that it wrote up the false recognitions — too many false alarms is more annoying than not recognizing at all. If you're thinking of installing one, set one rule first: alerts go to only one person, don't have the whole family buzzing; faces it can't recognize, have it just log them, not keep asking.

Editor Xiao He says | I had one at my door, and every time the delivery guy came it pushed a "stranger" alert, and a cat passing by at midnight pushed one too, so I turned push off in the end. This piece taught me that when choosing, first check how many false alarms, then check whether it can split alerts by person.

Source: The Verge (2026-09-25)

==

Leaderboard Flash

Third-party Intelligence Index · data as of 2026-09-23

1. Claude Opus 5.5 (top tier) | Anthropic | 58 / $5.98

2. Claude Fable 5.1 (top tier) | Anthropic | 53 / $7.63

3. GPT-6 Astra (top tier) | OpenAI | 53 / $3.26

4. GPT-6 Sol (top tier) | OpenAI | 48 / $1.06

5. Muse Spark 1.3 (top tier) | Meta | 48 / $1.60

First, why the switch: Arena's text and image leaderboards are still stuck on September 13, and the snapshot we got this time is the same one as last issue, no new data available, so here we've switched to a third-party outfit's composite index. It doesn't look at votes; it runs its own batch of questions to produce scores, higher is stronger. This outfit's scores got a fresh round on Claude Opus 5.5's launch day (September 23): Opus 5.5 took first with 58, Fable 5.1 and GPT-6 Astra both at 53. Fifth place Meta's Muse Spark 1.3 trails by 10. The highest-scoring open-source model is Xiaomi's MiMo-V2.6-Pro at 46. Within the same tier, who's cheaper — look at the last column: GPT-6 Sol runs about $1.06 per question, less than a seventh of Fable 5.1.

Blind Coding Leaderboard · not yet updated, data as of 2026-09-23

1. claude-opus-5.5-max | Anthropic | 1818 (1,219 matches, ±21)

2. gpt-6-astra-max | OpenAI | 1792 (4,325 matches, ±12)

3. claude-fable-5.1-max | Anthropic | 1755 (4,916 matches, ±11)

4. claude-opus-5-max | Anthropic | 1692 (14,351 matches, ±7)

5. gpt-6-sol-max | OpenAI | 1686 (1,521 matches, ±17)

This leaderboard tests coding: the same problem is given to two unnamed models, and people vote on who wrote it better, more votes means higher score. The snapshot we got this time is the same as last issue, not a single rank or score changed. Top is still Claude Opus 5.5 top tier at 1818, 26 ahead of second; but it's only played 1,219 matches with a ±21 margin, its position isn't solid. Second place GPT-6 Astra has played 4,325 matches with only ±12. Lots of matches, small margin — that's when a score holds up. When picking a model, look at the match count in parentheses first.

Agentic Ability Leaderboard · not yet updated, data as of 2026-09-24

1. Claude Fable 5.1 (Max) | Anthropic | 13.4 / 17.0 / 11.6

2. GPT 6 Astra (Max) | OpenAI | 11.1 / 13.7 / 6.7

3. Claude Opus 5 (High) | Anthropic | 9.8 / 7.5 / 11.9

4. Claude Opus 5 (Max) | Anthropic | 9.5 / 10.6 / 12.6

5. Claude Fable 5 (High) | Anthropic | 8.4 / 3.5 / 8.5

This leaderboard tests AI doing work on its own; the three scores are: how much better the work is than before, the rate tasks actually get done, and whether it can recover on its own when it hits an error mid-way. It's the newest of the four Arena leaderboards, refreshed to September 24, but it's still the same data as last issue, top five unchanged. One detail worth watching: fourth place Claude Opus 5 top tier scores 12.6 on "self-recovery after error," higher than first place, but only 10.6 on "confirmed done," meaning it can climb back from errors, but there are also plenty of tasks it fumbles from the start.

==

Section Picks

1. Recent "AI out of control" incidents may be a problem with the test paper

A September 25 TechRadar op-ed; the author went back through several recent incidents described as "AI suddenly going out of control" and offers another explanation: it's more like the testing method has holes, not that the model really has a mind of its own. The article calls out the evaluation stage — how questions are written, under what conditions it's tested, who grades it.

Comment: When you see "AI out of control again," ask first: under what conditions was this measured, who wrote the questions. That one line can block out more than half the headlines.

Source: TechRadar (2026-09-25)

2. NSA spends billions a year testing AI, where it goes is classified

A September 25 Washington Sun report: according to classified estimates, the US National Security Agency (NSA) spends on the order of billions of dollars a year testing AI models, and neither the exact amount nor the test list is public.

Comment: One piece of news makes one thing clear: scoring models is already big money, and that money doesn't have to be shown to outsiders. Only the public part is usable.

Source: Washington Sun (2026-09-25)

3. Make AI hand in evidence when it's done, not just "fixed it"

Executor appeared on GitHub on September 26, adding a step to coding assistants: from an idea coming in, writing a proposal, listing a plan, making changes, reviewing, accepting — every step has to leave something behind. The problem the author wants to solve is plain — AI says it's fixed, but it isn't.

Comment: What's valuable is making it prove itself. The project just open-sourced, try it for a week on a small module you won't miss if it's lost.

Source: GitHub / Hacker News (2026-09-26)

4. When debugging AI assistants, what work can only be done by humans

A post that hit Hacker News on September 25, the title asks a question: when debugging AI assistants, what things must be done by a human. The author lays the question out in the open and waits for people to add to it one by one.

Comment: The use of these posts is making a list. Write down the "human-only" parts and you'll know how far a tool can go for you.

Source: Hacker News / Traser (2026-09-25)

5. Put a meter on AI spending, block it on the spot when it goes over

AtlasBurn launched September 25; what it does is lay out AI call costs in real time: how much spent, where it went, you can set a cap and block it from running further when exceeded, so a runaway assistant doesn't blow up your bill.

Comment: If you're bringing AI into a company, install this first, more useful than looking at the bill at month's end. It manages how much is spent, not how well the work is done.

Source: Hacker News / AtlasBurn (2026-09-25)

6. Build your own search endpoint at home, just for AI

Sifthound on GitHub September 25: a search endpoint you can host yourself, made for AI apps to fetch online material. Its API is compatible with that paid search service on the market, so switching over doesn't need big code changes.

Comment: Worth trying if you want to save on a search service fee and don't want to hand over your query logs. Host it yourself and you maintain it yourself.

Source: GitHub / Hacker News (2026-09-25)

7. Draw an accurate map of a huge codebase, show it to AI first

cargo-atlas on GitHub September 25: draws a call-relationship graph for large Rust projects, with accuracy matching the compiler itself. When an AI coding assistant needs to check "who calls this function" or "where is this method implemented," just ask it.

Comment: The bigger the project, the easier it is for AI to change the wrong place. Give it the map first, less hassle than explaining over and over. Currently only supports Rust.

Source: GitHub / Hacker News (2026-09-25)

8. A company says it raised sales 60% with an AI coding assistant

A customer case posted on OpenAI's site September 26: Proaction says after using Codex, sales rose 60%, saving over 75 work hours. The numbers are published by the vendor itself, not third-party tested.

Comment: Discount the numbers first; what you can copy is the process: which categories of repetitive work they handed to the assistant.

Source: OpenAI site (2026-09-26)

9. AI that helps edit resumes is barred from making up experience for you

A free resume tool on Hacker News September 26: the AI only makes your experience read smoother and clearer, with a hard rule against adding things you didn't do, no registration needed.

Comment: The biggest fear with resume editing is AI embellishing, and one interview question exposes it. This restriction is more practical than its features.

Source: Hacker News / Nokku (2026-09-26)

10. Merge people from two separately-shot photos into one

tancky.io on Hacker News September 25: picks people out of different photos and composites a natural-looking group shot, for commemorative photos and family portraits.

Comment: Before posting a composite photo, ask the people in it whether they agree — that step matters more than which tool you pick.

Source: Hacker News / Tancky (2026-09-25)

==

What Everyone's Reading

  • Someone tore down the newly released Opus 5.5 and GPT-6 Luna, saying you can see the shadow of Chinese researchers on these two models (TMTPost)
  • The Pentagon wants to spend $30 million building an AI lie detector, to screen employees and find people leaking information (MIT Technology Review)
  • Starting from the extra 800 yuan your phone costs: what exactly is the most expensive chip waiting for (Huxiu)
  • Bloomberg asked a bunch of bosses, most feel their company isn't ready to use AI yet (Bloomberg Tech)

Tomorrow's Watch

① Arena's text and image leaderboards are still stuck on September 13, and the snapshots we got for the coding and agentic leaderboards this time are the same as last issue; wait for a refresh to see whether this week's new models made the boards

② NerfWatch has been running the same questions daily since September 25; wait for it to build up a week or two of records, then see which one is really trending down

③ On "AI assistant breaks into government website," wait for the platform and evaluators to release the full process and remediation notes, then judge whether it's a model capability issue or permissions opened too wide

2 replies

?
Ctrl + Enter to reply
Fang An Fan Zi

That NerfWatch timeline is what customers would actually pay for — being able to reproduce when it started getting dumber is what makes procurement willing to sign.

Terminology Police

"The test methodology has holes" — this phrasing is a first for me. Before, I'd panic the moment I saw "AI out of control"; now I ask first: under what conditions was it tested.