ReviewRadar · 2026-09-25
Issue 058 · Preview
Today's picks are three things: whether the glasses upgrade is worth it, whether AI grading each other can be trusted, and whether it can catch itself after making a mistake. After that, three leaderboards and ten highlights.
==
Top Updates
1. Most of the extra money for the new glasses buys "hearing clearly" and "shooting steady"
On September 25, Engadget put Meta's new smart glasses side by side with the previous generation and compared them item by item. The new model changed call audio pickup, camera quality, and battery, and the price went up too. The article's conclusion is blunt: this is a small step upgrade, but the price hike isn't small.
Wearables expert · Akai says | Wear them all day and you'll get it: no matter how pretty the specs are, if they press on your nose bridge, the temples heat up, and the battery can't last half a day, they end up in a drawer. With the previous generation I wore them for two full weeks, and to take a photo I had to stare at the target and hold still for a second before pressing. If this generation really fixes audio pickup and battery life, taking calls on the commute would actually be practical. If you're thinking of buying, do one thing first: go to the store and wear them for a full half hour, then decide whether to pay that extra premium.
Editor Xiaohe says | I've always wanted a pair of glasses that can take photos, but I'm most afraid of buying them and having them gather dust. This comparison helped me settle on an acceptance test: first check whether the battery lasts a day, then check whether taking a photo requires you to hold still.
Source: Engadget (2026-09-25)
2. Ten AIs turn in homework on 15 real-world problems, and their own kind grades the papers
An experiment that hit Hacker News on September 24 threw 15 real-world problems at ten AI assistants. Each first came up with its own solution, then took someone else's write-up and scored it. The questions, each one's answers, and who gave whom how many points are all posted on the page. The original page lists the date as September 23.
Product expert · Azhe says | I always say one thing: only when scoring power is taken back from the vendors is the exam worth anything. This time even the graders aren't human—they're peers taking the same exam at the same time, so nobody can tip anyone off in advance. Note that there are only 15 questions and ten models, so the rankings will wobble. Treat it as a method first, not as authority.
Editor Xiaohe says | From now on, whenever I see "some company's AI takes first place," I'll first ask: who graded the papers, and can I see the questions? If I can't, I won't treat it as a conclusion.
Source: Hacker News / Fix the World (2026-09-24)
3. Someone gave AI a self-reflection question: can you catch the mistakes you've made yourself?
tapoo, open-sourced on GitHub on September 24, tests only one thing: after AI screws up once, can it find out on its own and fix it back? The author puts it bluntly—when vendors release models they only show their best scores, and nobody writes about how they perform after making a mistake.
Programming expert · Lao Xu says | I buy this approach. What we fear most at work isn't that it makes a mistake once, it's that it keeps going after making one, and everything it hands in has to be checked from scratch. When picking tools I'll ask one more question: does it report errors, and can it roll back a step? The project just open-sourced and the question set is small, so try it for a week on a small module you won't mind losing.
Editor Xiaohe says | Last week I had AI fix a formula in a spreadsheet, and after it got it wrong it just kept calculating along with a straight face. Now I have it stop after making changes and show me before it touches anything.
Source: GitHub / Hacker News (2026-09-24)
==
Leaderboard Snapshot
Overall Intelligence Board · data as of 2026-09-25 (Artificial Analysis, live-updating)
1. Claude Opus 5.5 (highest tier) | Anthropic | 58
2. Claude Opus 5.5 (second-highest tier) | Anthropic | 56
3. Claude Opus 5.5 (mid tier) | Anthropic | 54
4. Claude Fable 5.1 (highest tier) | Anthropic | 53
5. GPT-6 Astra (highest tier) | OpenAI | 53
This board combines knowledge, problem-solving, coding, and hands-on tasks into one total score; the higher the score, the more well-rounded. Four of the top five are Anthropic, with the only one squeezing in being OpenAI's GPT-6 Astra. The same model is ranked by tier according to "how long it thinks"; the tier that thinks longer scores higher. The strongest in the open-source camp is Xiaomi MiMo-V2.6-Pro at 46; after that Zhipu GLM-5.3 at 45 and Moonshot AI Kimi K3 at 44. This board updates live, so the numbers today and tomorrow may differ.
Speed and Cost Board · data as of 2026-09-25 (Artificial Analysis, live-updating)
1. Fastest output: Celeris-1, about 1505 characters per second
2. Close behind: Mercury 2.5 / Mercury 2, about 770 / 761 characters per second
3. Cheapest for one job: GPT-6 Luna (low tier), about $0.0045
4. Second and third cheapest: Granite 4.2 3B / Ministral 3 3B, about $0.01
Same day, same company's data, but a different angle makes it completely different: the ones ranked first on the overall board don't make this board at all. The two fastest at output are built specifically for "instant replies"; the three cheapest are all small models—limited ability, but cheap. Pick tools to match your own work: if you use it once in a while, look at unit price; if you use it daily, look at speed and the monthly bill.
Human Blind-Vote Text Board · board not yet updated, data as of 2026-09-13
1. claude-fable-5-high | Anthropic | 1506 (30,057 votes)
2. claude-opus-4-6-high | Anthropic | 1505 (71,993 votes)
3. claude-opus-4-7-high | Anthropic | 1502 (60,002 votes)
4. muse-spark-1.2 (xHigh) | Meta | 1500 (3,227 votes)
5. claude-fable-5.1-max | Anthropic | 1498 (5,783 votes)
This board's scores come from ordinary people voting one-on-one: the same question is given to two anonymous models, and people vote on who answered better; more votes means a higher score. Four of the top five are Anthropic; Meta's muse-spark-1.2 is fourth, but it has only 3,227 votes and trails third place by just 2 points, so its spot is still shaky. The snapshot is still stuck at September 13, unchanged for two weeks, and none of the new models launched this week made it in.
==
Section Highlights
1. A bunch of assistants divide the work and stitch a long film segment by segment into a complete story
On September 25, Google Research's blog released a method: have multiple AI assistants divide the work, generate footage segment by segment, then stitch it into one complete long video, with the goal that the plot connects from start to finish without breaking immersion.
Comment: People who want to use AI to make long content should take note: what can be generated right now is mostly short clips, and the hard part has always been "making it still feel like one thing when stitched together."
Source: Google Research (2026-09-25)
2. Code written by AI, handed to another AI as a tester
Canary, launched September 25, turned "acceptance" into a product: after an AI assistant finishes writing code, it runs the program and specifically pokes at the places this change might have broken, and if it finds problems, hands them back to a human.
Comment: The approach is worth copying, and the boundary is clear too: it's only responsible for finding faults; whether to fix them and whether to ship is still up to humans.
Source: Hacker News / Canary (2026-09-25)
3. A poker bot in a 1.5MB package, running on your own computer
QuantPlay on GitHub on September 24: a set of poker programs written with AI approaches, with the whole engine about 1.5MB. Download it once and it runs locally—no sign-up, no server connection.
Comment: If you want to see what level a small AI-written program can reach, installing one and playing two hands is the fastest way. Don't use its performance to infer the level of large models.
Source: GitHub / Hacker News (2026-09-24)
4. Voice recorder pens on the market sell for $159; someone built one for $25
ZephClick on Hacker News on September 25: the author thought this kind of small hardware was too expensive, so he built a $25 alternative himself. Organizing and translating run on his own device, and it can also connect to his own account.
Comment: People willing to get their hands dirty can save some money; people who don't want the hassle can also see clearly how much of this small hardware's price is the shell and the brand.
Source: Hacker News / ZephClick (2026-09-25)
5. On the matter of arithmetic, someone made a calculator that only runs on your own computer
Axiom on Hacker News on September 24: a local-first AI calculator. Problems are computed only on your own device, the process is written out step by step, and nothing is sent to the cloud.
Comment: Using it to do accounts and check numbers is more reliable than asking a question in a chat box, and the key is that the process is visible. Being good at arithmetic doesn't mean it understands your business, so important numbers still need to be double-checked yourself.
Source: Hacker News / Axiom (2026-09-24)
6. Turning papers into tools you can click and try
IEEE Spectrum introduced Paper2agent on September 24: it automatically packages the method described in a research paper into a tool you can call directly, so others can try it without rewriting the code.
Comment: It saves researchers effort. Ordinary people can also get a bit of it: the capabilities in papers are turning into buttons you can click and try.
Source: IEEE Spectrum (2026-09-25)
7. Doing health checks on critical facilities like water, power, and healthcare; AI scans first for exposed doors
An action released jointly by Wiz and Google DeepMind on September 24: use AI to help scan the external entry points of critical facilities, list the "doors that aren't locked properly," and hand them to humans to deal with.
Comment: The direction is right, and it reminds us of one thing: what's scanned out is only a lead; who fixes the locks and which door to fix first still has to be decided by humans.
Source: Wiz blog / DeepMind (2026-09-24)
8. A security team hired an AI intern specifically to bang on its own code
GitHub Security Lab released a new tool on September 25: have AI automatically feed all kinds of messy input to its own code to see if it can trigger problems, and every hit is logged.
Comment: It's aimed at "letting AI find vulnerabilities for people first," and suits teams with a pile of code to protect. Before going live, try it on unimportant modules first.
Source: GitHub blog (2026-09-25)
9. The chat box isn't omnipotent: when to switch interfaces to let AI do the work
An article on GitHub's official blog on September 25 discusses several common scenarios for using AI to write code: having it change a small piece of code, the chat box is fine; having it touch many files at once, the chat box instead makes it hard to see what changed.
Comment: People who use AI to work every day should read it once; it can save you the detour of "one sentence in, whole page in chaos."
Source: GitHub blog (2026-09-25)
10. Someone sketched out "AI teaching itself," with Liang Wenfeng as the last author
On September 24, TMTPost interpreted a paper from DeepSeek and Tsinghua University. On the surface the article is about a sandbox for AI assistants, that is, an isolated environment that runs locked up; after section five it discusses the route of letting AI build its own environment and then use that environment to train itself.
Comment: For now it's still a research idea, and ordinary people don't need to dig into the details. One sentence is enough to remember: if this path really works, AI will progress faster, and it will need people watching to keep it from going off track even more.
Source: TMTPost (2026-09-24)
==
What Everyone's Watching
- The CEO of security company Darktrace says AI assistants that can act on their own are becoming a "new kind of insider" in companies (Bloomberg Tech)
- A company turned off all AI permissions for new hires, letting people get hands-on first before gradually opening them up (Hacker News / Valon)
- McDonald's AI drive-thru order-taker Archy is live, able to take orders in English and Spanish (Bloomberg Tech)
- Revolut is piloting face-scan checkout in the UK, aiming to cut one payment step (Engadget)
Tomorrow's Watch
① Arena's blind-vote boards are still stuck at the September 13 to 15 snapshots; wait for them to refresh and see whether new models like Claude Opus 5.5 and GPT-6 Astra make it in
② Ray-Ban Meta Gen 3 just went on sale; wait for people who've worn them a full week to report back: how heavy are the temples, how many hours can calls last
③ tapoo's "self-healing after errors" exam is self-set and self-scored; wait for others to rerun the same set of questions and see whether the rankings match up
Physix Frontier