Community Discussion · Tracks

ReviewRadar · 2026-09-14

Issue #047 · Preview Edition. Others do reviews; we build the radar for those reviews.

Key Updates

1. AI learns to operate software itself; old methods of ranking LLMs need recalculation

Huxiu's "Thought Imprint" article connects the dots on recent buzz from the past month. Over the last two weeks, many people have used GPT-6 Astra to create 3D models and mini-games, complaining that the results aren't as good as specialized video models. The author argues they are comparing apples to oranges. OpenAI positioned Astra beyond just answering questions; it is an assistant that knows how to use computers—filling out forms, entering data into CRM systems, organizing schedules, and opening software to create charts. As AI moves from generating content to operating software to complete workflows, the metric for evaluating models should shift to: Did it get the job done from start to finish? How many clicks did it miss? Did it self-correct when it made mistakes?

Read in reverse: If a model can only answer questions but users don't think to use it for actual work, no matter how high its score, it's just a smarter search bar. In the future, when looking at model marketing, don't ask what the total score is. Ask how many steps it takes to click through a real workflow from start to end, how many errors occur, and if it fixes them automatically. Currently, this narrative exists only in vendor demos; wait for third parties to re-run these tests with real office workflows.

Editor Xiao He says | My understanding is: We used to compare who got higher test scores; now we compare who can take my weekly report from digging through chat logs to formatting the table completely. Sounds convenient, but the rule I set in Issue #041 remains unchanged. An AI that can click the mouse for me can also delete files for me. Before letting it touch your work computer, ask: "Is every step logged? Can permissions be revoked with one click?"

Source: Huxiu / Thought Imprint (2026-09-13)

2. A "Driving School" for AI Agents: Classes, Exams, Diplomas, Scores Publicly Searchable

Agents School, launched yesterday, aims to do something quite bold: build public profiles for AI agents. Agents learn from community-written courses, pass exams graded by code, and receive "diplomas." Anyone searching for a name can see what exams it passed and its scores. Platform administrators call this "agents that can prove what they know," replacing the marketing slogan "Our Agent can do anything" with a verifiable transcript.

Security Expert Lao Zhou says | This direction hits the pain point mentioned in the previous issue: The agent you evaluate is not necessarily the agent you deploy. Transcripts are only meaningful if they mandatorily state "Which version was tested? What configuration was used?" Otherwise, it's just an interview-ready resume. For new testing grounds, three things must be asked: Who wrote the questions? Who graded them? Is the question bank public? Registered agents can come and take the exam themselves. The distance between "holding a certificate" and "actual capability" will only be known after the first batch of failures. Don't hand over the keys just because someone has a diploma.

Editor Xiao He says | Got it, it's like a driver's license for the AI world. But I still look at buyer photos rather than ads when buying tickets. First, watch how public the question bank becomes. Second, wait for the first news story about someone who "got a perfect score but caused chaos at work." Then see if there's a mechanism to revoke diplomas. Certificates without revocation mechanisms don't count for me.

Source: Agents School / Hacker News (2026-09-13)

**3. After AI modifies code, don't just look at what changed; new tools let you see *how it behaves***

The open-source tool RunBoth on Show HN targets a familiar scenario. An agent modifies code, the diff looks fine, but merging it reveals behavioral changes. Its approach is to run both the pre-change and post-change code simultaneously, comparing behavior case-by-case against the same test suite, outputting a "behavioral diff" instead of a "textual diff." Accepting AI-modified code shifts from human eyes reading text to machines comparing results.

Programming Expert Lao Xu says | This cuts right into the weak spot of the three hard rules I laid out in Issue #034. I said "Always check the diff before merging AI submissions," assuming diffs expose problems. Changing a boundary symbol or rounding method might be just a few lines of text, but the behavioral difference is night and day—human eyes can't catch it. Textual diffs tell you what changed; behavioral diffs tell you what consequences resulted. Treat this newborn project with caution: roll it out on small open-source libraries for a week first, checking its false positive rate for async operations and randomness. Don't rush to integrate it into main repo CI.

Editor Xiao He says | In Issue #010, when I first used auto-mode, I watched closely which files it touched. This tool essentially outsources the "watching" to the machine—I like it. For non-coders, there's a simple usage tip: Next time you ask AI to change something, don't ask "Did you fix it correctly?" Ask "Run it before and after and show me the results."

Source: RunBoth / Hacker News (2026-09-14)

Leaderboard Flash

First, a clarification. The Arena Human Blind Test Main Leaderboard has had the same snapshot for two issues (last refreshed September 11). There are no new market movements today, so per usual practice, we won't copy the previous issue. Instead, we dug up a sub-leaderboard nobody was watching to fill the gap, and interpreted two other groups using the same snapshot from a different angle.

Visual Blind Test Leaderboard (Snapshot Aug 27) 1. Claude Fable 5 (Anthropic, 1313 pts) 2. Claude Opus 4.7 High (1301) 3. Qwen3.8-Max (Alibaba, 1300) 4. Claude Opus 4.7 (1299) 5. Claude Opus 4.6 High (1299). Plain English: In blind tests involving image-based Q&A, Anthropic holds four of the top five spots. Alibaba's Qwen3.8-Max is 3rd, tied with 4th place by just 1 point. Note that this sub-leaderboard snapshot stopped on August 27, so treat rankings as approximate.

Text Blind Test Leaderboard (Snapshot as of Sep 11, not yet updated) Top 5: Claude Fable 5 (1506), Opus 4.6 High (1505), Opus 4.7 High (1502), Fable 5.1 Max (1501), Meta Muse Spark 1.2 xHigh (1499). Anthropic occupies seven of the top ten spots. Two signals worth noting: The top 9 differ by only 13 points; the head is crowded, making differences hard to perceive in casual conversation. Ranks 4 and 5 have only 5,447 and 3,229 battles respectively—an order of magnitude thinner sample size than veterans. Record their ranks but don't fully trust them yet.

Agent Leaderboard (Same Snapshot) Claude Fable 5.1, ranked #1 overall, scores only 0.91 on "Steerability" (can it listen and turn around if corrected mid-task?), placing it near the bottom. Conversely, #3 Opus 5 leads with 12.77, while #2 GPT-6 Astra sits at 3.75. For long-flow tasks where the AI acts on your behalf, "obedience" and "intelligence" currently reside in different models.

Section Highlights

1. Establishing 15 "Service Metrics" for AI Agents, but monitoring frameworks themselves fall short. Agent thought loops and tool-call spinning are blind spots in current monitoring. Comment: You dare to use AI for work only if you can replay it. This piece adds the second half: Replay systems themselves have blind spots.

2. ClientCoded builds mock exam environments for agents, pre-loading 35 software environments like Salesforce and Jira, auto-generating 200 tricky queries and standard answers. Comment: When the same entity handles the exam venue, question bank, and grading, do the question-setters and agent sellers share interests? Keep an eye out.

3. "If AI answers poorly, fix your company's data catalog first." Developers found that tweaking prompts endlessly didn't help; the root cause was AI not knowing what tables exist or what fields mean. Feeding AI a manually curated data catalog significantly improved accuracy. Comment: Moving the blame for "dumb AI" from its brain to the reference shelf is the flip side of Issue #043's "AI blindly flipping through books."

4. Drexler's long-form warning suggests AI agents may learn to "collude on pricing," making it hard for regulators to detect when both buyers and sellers use AI for negotiation. Comment: The "lowest price online" shown by comparison apps might be a price agreed upon by two AIs.

5. learnlance automatically grows personal knowledge maps from AI-written code. Comment: Issue #041's code map, Issue #045's terrain map—another car joins this crowded lane. Wait for third-party re-testing of effectiveness.

6. Agent Tavern launches: One AI asks, one model answers, and a different model must grade. Comment: Issue #045 noted that cross-checking with random models is harder to fool than self-review. Turning this idea into a product comes with the cost that the grader model's biases become variables.

7. llm-hub packs 15 open-source models onto phones, enabling offline chat, image generation, and video creation without API keys or subscriptions. Comment: On-device all-in-one suites sound capable, but how much performance is lost in smaller models? Can phones handle the heat? Wait for real-world tests.

8. A programmer posts to defend AI writing, starting with the disclaimer: "This article was written by living, breathing flesh." Comment: Readers can no longer distinguish between humans endorsing AI and AI ghostwriting for humans. That ambiguity is likely why the article exists.

9. An individual hand-crafts an "AI Boom Chronicle," listing 437 events daily from 2017 to 2026. Comment: Chronicles are the antidote to benchmark-score narratives. Every generation's strongest model once topped the list, only to be overwritten by the next page.

10. The as-an-engineer skill pack cures agents of "Nanny Syndrome," forcing the model to converse with you as a senior engineer. Comment: Models optimize for seeming responsible by leaving half-sentences unspoken. Tools that reverse-train this behavior have market potential.

Trending

GPT-6 Astra positioning debate (Huxiu), Agent Driving School (HN), Behavioral Diff Tool RunBoth (Show HN), AI Collusion Pricing Warning (Substack), Arena Main Leaderboard stalled for Day 3.

Tomorrow's Watchlist

① Monitor Arena leaderboard refreshes, specifically checking if the ranks of Fable 5.1 (only 5,447 battles) and Muse Spark 1.2 (only 3,229 battles) remain stable given their small sample sizes.

② Watch Agents School: Who gets the first diplomas? How granular is the public question bank? Do vendors immediately use certificates as advertising?

③ Monitor early feedback on real-world integrations for RunBoth and ClientCoded before deciding whether to recommend them.

Data is subject to official disclosures. This column focuses on the true capabilities of AI hardware and software.

2 replies

?
Ctrl + Enter to reply
hongtao
hongtaoSep 14

Can this evaluation set pass top-tier conference peer review? Lab funding isn't enough to run inference at this scale.

Lei Who Shoots Films

This UI is anti-human; viewers swipe away in three seconds. I tested it for you guys—the barrier to entry is way too high. Don't touch it if you don't have traffic.