Community Discussion · Tracks

ReviewRadar · 2026-09-05 (Issue 38 · Preview)

Physical World Frontier Review · 2026-09-05 (Issue #038 · Preview)

Today's main leaderboard switched sources. The snapshot for the Arena Human Blind Test Leaderboard stopped updating on September 3rd and hasn't refreshed, so following our usual rule of not copying the previous issue, we are using LiveBench and Artificial Analysis data scraped in real-time on September 5th as the main leaderboard, with Arena data serving only as a reference.

I. The "Practical Exam" for AI Coding is Here: AWS Turns Real Cloud Service Tasks into Test Papers

On September 4th, Amazon open-sourced a set of test questions called AWS-bench, specifically designed to test how well AI assistants that can write code, modify configurations, and run commands perform in real AWS cloud environments. Previous coding exams mostly involved modifying small snippets of programs; this time, they directly brought tasks like cloud deployment, troubleshooting, and configuration fixes—things that happen daily in companies—into the exam hall, and also allowed vendors to package their models and tools together for testing.

Programming Expert · Lao Xu says | The direction hits the nail on the head. No matter how impressive the official demos are, they can't beat a unified environment test paper. But remember the lesson from Huawei's comparative experiment on August 22nd: test scores are just scores, and the scope of questions and environment configurations are all defined by the organizers. Before taking on real projects, it's still the same advice: run it on your own codebase for a week first, don't rush into production environments.

Editor Xiao He says | From an ordinary user's perspective, when choosing AI coding tools in the future, besides looking at whose marketing copy looks good, you can wait for another "AWS Practical Exam Report Card." Continuing the old principle I mentioned on August 26th: new test papers should be noted but not entered immediately; wait until the rankings stabilize for one or two rounds before deciding.

II. Meta's New Model Hits Global Third Place Three Days After Launch, Price is Only One-Sixth of First Place

LiveBench's real-time leaderboard today shows that Muse Spark 1.3, released by Meta on September 2nd, scored a total of 81.6, jumping straight to third place overall, trailing only Claude's two flagship models (83.4 and 83.0). What's even more striking is the cost: it averages $0.219 per successful task, while the number one ranked Claude Fable 5.1 costs $1.212, nearly six times higher. On Artificial Analysis's Intelligence Index leaderboard on the same day, it scored 62 points, entering the top tier, and its ultra-high-speed version remains the fastest output among the leaders, at 182 tokens per second (tokens are the character units used for AI pricing and output).

Product Expert · A Zhe says | In yesterday's issue, I said the battle for entry points has shifted from "who is strongest" to "who is most worth using." Meta is throwing punches with both fists here: proving capability on the leaderboards and flipping the table on price. Last round, their Muse Spark 1.2 made headlines with a data-sharing discount plan offering 1.3-8% off list price; this time, the new version directly slots into third place among the top tier. The key point is whether other vendors will respond to the pressure of three leaderboards simultaneously driving down prices. As usual, newly ranked positions tend to fluctuate, so note it down first.

Editor Xiao He says | With such a low price, my first reaction was to try writing weekly reports with it. But following the rule I set on August 28th, no matter how cheap it is, someone needs to empirically test that it's "not stupid" before I recommend it. How Meta's AI uses data and what account permissions are granted are still unclear to ordinary people, so let's wait for third-party reviews first.

III. Will Bosses Accept AI-Written Code? Empirical Ratio: Claude 84%, Humans 85%

An empirical study posted on arXiv pulled code modifications submitted by various AI programming agents and calculated the "merge rate," i.e., how many proposed changes were actually accepted by project maintainers and merged into the codebase. The results were quite surprising: Claude-series agents had an 84% merge rate, OpenAI's Codex had 74%, fully automated Devin had only 43%, while human contributors had 85%. AI-written code is now almost as "review-passable" as human code, but the gap between different tools is larger than the gap between humans and AI.

Programming Expert · Lao Xu says | Merge rate is an honest metric; it measures whether other people's projects accept the code, not self-reported benchmark scores from vendors. I said on August 22nd that "as AI output increases, verification must keep up," and this study completes the other half: AI that keeps up with verification produces quality that truly matches humans. However, Devin's 43% needs to be interpreted correctly; fully automated tools often take on dirtier, harder tasks, so the denominators are different, and you can't simply say it's twice as bad as Claude.

Editor Xiao He says | The most useful part of this for people like me is that in the future, when hearing claims like "AI already writes better code than programmers" or "AI code is unusable," you can ask back, "What's the merge rate, and what project was it tested on?" Continuing my old saying from August 29th: verify AI-reported results with humans before drawing conclusions.

Leaderboard Flash

LiveBench (Third-party comprehensive exam, dynamic question rotation to prevent memorization, 09-05 Real-time)

# Model Total Score Autonomous Coding Cost Per Task
1 Claude Fable 5.1 Max Effort 83.4 66.1 $1.212
2 Claude Fable 5 Max Effort 83.0 62.2 $1.439
3 Muse Spark 1.3 xHigh (Meta) 81.6 64.1 $0.219
4 GPT-5.6 Sol Max Effort (OpenAI) 81.0 56.2 $0.515
7 Kimi K3 (Top Open Source) 79.2 62.2 $0.348
14 DeepSeek V4 Pro 0813 (Open Source) 77.4 54.9 $0.044

Plain English | This leaderboard rotates questions irregularly, leaving little room for AI to score high by memorizing the question bank, so the industry treats it as a "physical exam." The total score difference among the top six is only 2.4 points, making it very tight. The real watershed is in the cost column: DeepSeek V4 Pro costs about 4 cents per successful task, the lowest overall, while Kimi K3 is the top-ranked open-source model. "Autonomous Coding" tests the ability of AI to independently modify codebases and run tests, scoring generally ten-plus points lower than the total score, indicating that current AI still has shortcomings when acting as "self-working assistants."

Artificial Analysis Intelligence Index (Multi-subject comprehensive exam with transparent pricing, 09-05 Real-time)

# Model Intelligence Index Cost Per Task Output Speed
1 Claude Fable 5.1 (max) 66 $3.69 66 tokens/sec
3 Claude Opus 5 (max) 63 $2.34 —
6 Muse Spark 1.3 (max) 62 — —
9 GPT-5.6 Sol (max) 61 $0.95 71 tokens/sec
11 Muse Spark 1.3 (xhigh) 61 $0.55 182 tokens/sec
12 Kimi K3 (max) 60 $0.84 —

Plain English | This organization synthesizes a bunch of public exams in math, reasoning, coding, etc., into a single index, while also marking the actual cost for each model to "complete a job," essentially providing a health check report with prices attached. The first place is only 5 points higher than ninth place, but costs nearly 4 times more ($3.69 vs $0.95); Muse Spark 1.3 high-speed version outputs 182 tokens per second, the fastest among the top tier. When scores are close, the competition becomes about who is cheaper and faster.

Arena Human Blind Test Leaderboard (Not updated on 09-05, data as of 09-03, for reference only)

# Model Vendor Blind Test Elo Score
1 claude-fable-5 Anthropic 1507
2 claude-opus-4-6-high Anthropic 1505
3 claude-fable-5.1-max Anthropic 1504

Plain English | Arena's method involves real humans chatting with two anonymous models simultaneously and voting; higher Elo scores indicate greater popularity. This snapshot is identical to the previous issue, so per our rules, it is not treated as the main leaderboard. The top three are still monopolized by Anthropic, with only a 5-point difference between 1st and 4th place, which constitutes close combat in the human voting leaderboard.

Section Highlights

1. A skill that makes "AI speak short sentences" goes viral, claiming 42% token savings. A developer created an optional response style called Macha, prompting AI to speak in concise sentence structures typical of South Indian English, claiming an average saving of 42% of tokens (the smallest billing unit for AI; fewer words mean lower bills). The money-saving idea is as simple as saving fuel in cars, but whether speaking 42% less will also omit critical information has not yet been verified through controlled experiments.

2. Renting a "workstation" Linux host for AI. Developers open-sourced Kovavue: any agent that speaks the MCP protocol (a universal docking language between AI tools) can rent a dedicated Linux machine managed by the provider, operating on it like a real person, with accounts handled via signature custody. Instead of letting AI borrow your computer to work, give it a rented machine that won't hurt if it breaks; permission isolation is the ticket for agents to enter production.

3. Open-source software sets "traps" for AI crawlers. Maintainers of NetworkManager, a Linux network management component, launched a mechanism hiding "canary" markers visible only to programs within documentation; if AI agents scrape content without adhering to authorization policies, citing them will expose the violation. This turns "whether AI follows the rules" into detectable evidence, as the open-source community begins building its own audit tools.

4. Montgomery, a visual training toolbox written in Rust, is open-sourced, allowing object detection and image segmentation training on ordinary GPUs; it is experimental in nature. It's still far from ordinary users, but the fact that "visual model training is no longer exclusive to the Python ecosystem" is a positive signal for edge-side AI hardware.

5. Who decides "AGI is here"? OpenAI claims "The AGI era has arrived" alongside its new model, while The Verge podcast dissects each company's definitions point by point. There is still no universally accepted exam for AGI, so anyone can draw lines based on standards favorable to themselves. Treat "breakthroughs" without public exams as marketing first.

6. AI coding modifies code, but git can't explain why. A new tool called Casefile records the context of every AI code modification, allowing code history to explain commits involving AI participation. As AI output increases, "who changed it and why" has become a new pain point for teams.

7. Empirical test of Google Slides' AI beautification feature. The author used it to beautify slides, resulting in the entire page being returned as a non-editable image, impossible to modify. The author has a competitive stance against Google, so discount the conclusion, but the pitfall of "AI features turning editable objects into dead images" is real; back up important files before clicking beautify.

8. Giving AI philosophy exams. A developer initiated an AI philosophy competition, arguing that while large models are getting stronger in math and coding, subjects like philosophy, which have no standard answers, lack test creators. Philosophy exams inevitably carry subjectivity, so scoring rules are more important than scores, but the next battlefield for AI evaluation is indeed questions that "cannot be automatically graded."

9. Hours of training enable humans to recognize AI-generated faces. An experiment at the University of Southampton showed that after brief training, participants' ability to identify AI faces improved significantly, whereas ordinary people's recognition rate for AI face-swapping was previously close to guessing. AI face forgery evolves much faster than detector updates; rather than relying on platforms, spend an hour training your naked eye.

10. Reverse AI: AntiAgent makes machines ask questions, humans think. This product reverses the relationship of "humans giving instructions to AI": AI is responsible for asking sharp follow-up questions, and humans are responsible for thinking and answering, targeting scenarios where "procrastination often stems from unclear thinking." In an era where everyone practices prompt engineering, some are starting to doubt the direction; regardless of whether this experiment succeeds, it's worth a look.

Everyone is Watching

Anthropic finalized a $15 billion credit line to pave the way for IPO, pushing external trustee systems under a $2 trillion valuation into the spotlight; Moonshot AI is reported to be considering a Hong Kong listing as early as this year, planning to raise up to $5 billion; Altman admits that "curing cancer" alone won't win over a public skeptical of AI; Tesla Cybercab faced an investigation filed by US federal safety regulators on its first day of operation in Austin.

Tomorrow's Focus

① Will the Arena snapshot update on 09-05? The top three in image-to-video are locked within 16 points, see who falls behind first.

② The substance of Meta Muse Spark 1.3's 3rd place on LiveBench: will the ranking fluctuate once battle samples accumulate, and is the bargain of $0.22 per task worth it?

③ Will reverse experiments like AntiAgent, where "AI asks and humans think," attract a second wave of imitators? The reputation data on Show HN is worth monitoring.

Complete leaderboard tables and dual-commentary versions are available on the official review page.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts