ReviewRadar · 2026-10-02
Issue 065 is here. Three highlights this issue: one about how much AI relies on humans pointing the way when finding bugs in code, one about an AI built by a three-person Japanese team taking first place in a vulnerability contest, and one about AI learning to solve problems it's never seen by watching cat photos.
Top Updates
1. If you don't tell the AI where the bug is, it basically can't find it
Someone made a new exam that throws an entire real code repository at an AI and just says "there's a bug in there, find it yourself" — no mention of which file or which line. Turns out most AIs handed in a blank sheet. Old Xu's take is blunt: in real projects nobody marks the location of the problem for you, and being able to find it yourself is what counts as skill; if you want to try it, start with a small repo you won't cry over losing and roll with it for a week, don't go straight at the one you're maintaining. Xiao He says from now on when she has AI check things, she'll first spell out "roughly where it's wrong."
2. An AI built by three people in Japan took first place in a bug-finding contest
Japanese small company Layer8 has only three people, and their bug-finding AI is called Trident, which ranked first on the bug bounty platform HackerOne for Q3 2026. Old Zhou warns that an AI that can find vulnerabilities can also be turned around and used against you — first ask three things: who authorized it to scan, what did it do along the way, and can you one-click stop it if something goes wrong. Xiao He says this makes her even more afraid to casually grant permissions to company systems, and she'll get around to canceling the assistants she doesn't use often this week.
3. By looking at cat photos, AI learns to solve problems it's never seen
He Kaiming's team's new research: first let the AI look at lots of cat photos, then have it take a test specifically designed for "unseen new problems," called ARC, to see if it can grasp the patterns in problems it hasn't been taught. This test doesn't allow memorizing answers — it tests whether you can draw inferences. A-Zhe says AI is mostly good at memorizing answers and gets stage fright on new problems; whoever truly learns to draw inferences first is the one you'd dare hand more autonomous work to.
Leaderboard Brief
This issue's three leaderboards (human blind-vote text board, coding board, hands-on capability board) are the same as last issue, no updates, data as of 2026-09-30, so don't use the rankings to pick tools just yet. If you want to pick an AI for writing code, first check whether the number of matches in parentheses is enough.
Section Picks (10 items)
- A model that can only pick options from a list will never learn to talk: locking speech down to only picking options saves money, but it won't buy you conversation.
- Roundup: what level is the AI that makes robots move at now: robots are still far from entering homes, so first see clearly what it can and can't do.
- Why large models didn't become parrots that only recite: grasping patterns is real skill, and making things up is a real flaw — you have to guard against both.
- Every step was approved, yet the whole thing still went wrong: managing things can't just look at whether each action is compliant, you have to look back at whether the whole thing got done.
- Cloudflare open-sourced a batch of small models "for making judgments": not every judgment needs a large model to show up; a small model is enough and saves money.
- Nvidia released a toolkit specifically for locking up "out-of-control AI": giving AI a one-click lockdown switch is more practical than assigning blame after the fact.
- Give coding AI an independent little room where they don't disturb each other: when multiple AIs work in parallel, isolate them first so trouble doesn't spread.
- Record the work AI has done, and next time just replay it: no need to spend money again every time on repetitive work.
- One small command helps several AIs each work on their own part of the same code: split the branches first, it's less hassle than hunting for conflicts afterward.
- Someone open-sourced the whole toolbox for "training large models": the more public the recipe for training large models, the easier it is for newcomers to get started.
Everyone's Watching
Google Gemini 4 Argon topped several leaderboards on the day it launched; a three-person Japanese team's AI took a quarterly first on a bug bounty platform; Nvidia released an open-source security tool that can one-click isolate out-of-control AI; a new exam proves that without narrowing the scope, AI can't find bugs in code; He Kaiming's team can get AI to learn unseen problems by watching cat photos.
Tomorrow's Watch
① The leaderboards didn't update this issue, so watch whether Arena recalculates in the next few days, or whether a new board picks up.
② For that new exam where "without narrowing the scope it can't find the bug," watch whether anyone reruns it with other AIs.
③ AI took first place in bug-finding, so watch whether this toolset will be required to disclose authorization and keep a full trail.
Physix Frontier