Community Discussion · Tracks

ReviewRadar · 2026-09-02

Physical World Frontier Review · 2026-09-02

Physical World Frontier Review · 2026-09-02 · Issue No. 035 · Preview Edition

5 minutes a day to understand which AI tools are worth using for you. Today's three big stories: OpenAI admits its new model is too good at "hacking," Huawei's open-source framework gives AI coding tools a lesson in "hands and feet matter more than the brain," and robot vacuums get a national unified exam.

I. OpenAI Admits New AI Is Too Good at "Hacking," Safety Rating Raised to Highest Level

On September 1, OpenAI announced that its upcoming new model, Astra, is the first product to cross its internal "Critical" (highest danger level) cybersecurity threshold. According to this rating, it no longer needs humans to teach it step-by-step; it can find vulnerabilities and write attack plans on its own. OpenAI also stated that this capability will not be open to general users, but only selectively released to a few trusted defenders. Remember the issue from August 20 where we followed up on the "unreleased model crossing boundaries on Hugging Face" scandal? After that incident, OpenAI delayed the development of another unreleased model, while Astra actually received more attention.

Cybersecurity Expert · Old Zhou says| Capabilities will eventually emerge; when they are released, to whom, and whether there is an audit—that is the real substance of this news. When a model can autonomously complete attack actions, guardrails must shift from the "prompt layer" down to the "environment layer," locking down the systems it can touch and keeping full behavior logs. Also, a reminder: "Critical" is OpenAI's own rating standard. Whether the rating is credible or not, following our usual rule, wait for third-party verification.

Editor Xiao He says| It sounds scary, but ordinary people neither need nor should have access to it. What I care about is the flip side: AI that can find vulnerabilities can also help patch them. For ordinary users, the old advice remains unchanged: don't enable important account permissions unless necessary, and don't rush to authorize.

II. Same AI Brain, Different "Hands and Feet" for the Exam, Coding Score Increases by 3 Points

On September 1, Huawei's open-source AI agent platform openJiuwen released a technical report, setting new records on two AI coding "practical exams." The first is SWE-bench Verified, which selects 500 questions from real GitHub bugs to fix code; it scored 82.6%. The second is Terminal-Bench 2.1, simulating a real terminal environment for complex tasks; it scored 87.19%. Crucially, they conducted a "same-model control" test: using the same model as the top system on the leaderboard, the score was 3.4 points higher; switching to the same model used by Claude Code, it still surpassed it. To translate: the model determines the AI's foundational intelligence, but the execution framework wrapped around it determines how far it can actually go. In the future, choosing AI coding tools won't be enough by just asking "what model does it use"; you have to ask "how does it drive the model to work."

Programming Expert · Old Xu says| "Tools are amplifiers, not engines." In Issue 022, I called out shell-wrapping scams; this paper is a positive sample—public paper, open-source code, installable product. This posture is better than just releasing posters. But 82.6% is a self-reported score; third parties haven't re-run it yet. Following my old rule, discount it first and keep a record. Additionally, it still failed to handle 17 out of 100 questions smoothly. Before integrating it into your own projects, run real tasks for a week to see where it breaks.

Editor Xiao He says| If the same model can score a few more points just by changing the wrapper, then I really shouldn't just look at the brand when choosing tools. However, scores are just scores; I'll wait for third-party results running their own projects before deciding whether to switch.

III. Robot Vacuums Are Getting a "National Unified Exam," 11 Test Items Explained Clearly

CCTV News reported on September 1 that China recently led the formulation of the international standard "Performance Evaluation Methods for Household and Similar Robots," establishing a unified performance evaluation framework for the global household robot industry, covering 11 core tests including obstacle avoidance, slope operation, and energy consumption. Previously, parameters from various brands were all based on their own lab standards; nobody knew how claims like "covers 98% of the floor" were measured. With this unified exam, the numbers on promotional pages from different brands finally have the possibility of being comparable.

Wearables Expert · Akai says| In reviewing hardware over the years, what I fear most is "each one sets their own questions and grades themselves." Unifying standards is the first step; we also need to watch the second step: who tests, how they test, and whether reports are public. For ordinary users, when buying a robot vacuum in the future, you can at least ask, "Do you have scores for these 11 items?" Compliance costs for small brands will rise, and prices might increase first.

Editor Xiao He says| As someone whose robot vacuum bumps into slippers every day, I feel the most connection with the words "obstacle avoidance." If the phrase "smart obstacle avoidance" on promotional pages doesn't have test scores to back it up, I won't believe it anymore. This is very practical for ordinary people.

Leaderboard Flash Report · Who Was Strongest This Week

Arena Human Blind Test Three Leaderboards (Snapshot Aug 27–31, scraped Sep 1, same source as last issue so not refreshed yet; see below for new data)

  • Chat Leaderboard #1: Claude Fable 5 (Anthropic), 1507 points, 25,824 battles
  • Chat Leaderboard #4: Meta's Muse Spark 1.2 (xHigh), 1498 points, but only fought 3,247 battles
  • Code Leaderboard #1: Claude Opus 5 (Max) (Anthropic), 1688 points
  • Code Leaderboard #2: Kimi K3 (Max) (Moonshot AI, open source), 1674 points
  • Vision Leaderboard #3: Qwen 3.8 Max (Alibaba, open source), 1300 points

Plain language interpretation: Anthropic occupies four seats in the top five of the chat blind test. Meta's #4 spot has battle samples 5-6 times fewer than other models, so the confidence interval for the score is obviously wider; don't take the ranking seriously yet. Two of the top five in the code leaderboard are open-source models.

Coding Practical Exams (SWE-bench / Terminal-Bench, self-reported in openJiuwen technical report on Sep 1, no third-party reproduction yet)

  • SWE-bench Verified: 82.6%, 3.4 percentage points higher than the strongest system using the same model
  • Terminal-Bench 2.1: 87.19%, compared to Claude Code's 83.8%
  • Same-model control: Switching to the same model yields 84.04%, still surpassing Claude Code

Plain language interpretation: Using the same model with a different driver framework allows scoring 3+ points higher on both exams. The evaluation of "can AI do the work" is shifting from "looking only at the model" to looking at "model plus framework" together.

Artificial Analysis Intelligence Index (Scraped Sep 2, not used in previous issue)

  • #1: Claude Fable 5.1 (max), 66 points, cost per task $3.69
  • #8: GPT-5.6 Sol (max), 61 points, cost per task $0.95
  • #12: Kimi K3 (max), 60 points, cost per task $0.84

Plain language interpretation: The gap between #1 and #8 is only 5 points, but the price difference is nearly 4 times. Domestic open-source Kimi K3 and Zhipu GLM-5.3 have squeezed into the top tier, with costs kept under $1. "Cheap and generous" has collective presence on the report card for the first time.

Section Highlights

1. "Pulling the Plug Mid-Road" for AI Agents, Measured at 127 Milliseconds. An engineer's hands-on article shows that AI coding agents often hold your cloud keys and external network channels while running tasks. The author created an interception mechanism to cut off network access mid-task, taking only 127 milliseconds from decision to disconnection. Comment: "The environment decides" gains another deployable sample; guardrails don't necessarily slow things down.

2. New Method for Health Checks on "Exams That Test AI". BenchMIRT, released early morning on Sep 2 on the HuggingFace community, no longer just scores models but audits the evaluation question bank itself, checking what each question is actually testing and if models are "memorizing answers." Comment: Give the exam paper a health check first, then the scores have credibility.

3. Price List for 900+ AIs in One Real-Time Table. Show HN project Indextkn lists real-time pricing for 976 models from 17 vendors. As AI price hikes and cuts become more frequent, you don't need to visit each official website to compare prices anymore. Comment: The tool is free to check, but you need to sample-check if the data is complete.

4. A Bash Script Locks AI Coding Assistants in a "Quarantine Room". Open-source project Dev-sandbox uses one script to run AI agents in isolated containers, specifically curing the ailment of "AI working with the keys to your main machine." Comment: Don't rush to let AI touch your main machine; give it a room that can be demolished anytime first.

5. How Does AI Writing Detection Actually Work? Pangram founder revealed details on a podcast: the detector's principles, why it misjudges, and how to read the scores. Comment: Continuing our stance from Issue 028, detectors can only serve as hints, not verdicts.

6. Small Plugin Stops AI from Secretly Using Outdated Dependencies. Open-source tool yul automatically checks versions when AI writes new dependencies into a project, blocking outdated or insecure ones directly. Comment: As AI output volume increases, the verification step must be built into the assembly line.

7. Wasmer Releases New SDK, Providing Local "Safe Houses" for AI Agents. Code executed by agents is locked in independent environments, not touching the host file system. Comment: Another regular army joins the isolation track; this is good for ordinary users.

8. Public Archive for Traces Left by AI Agents After Work. Show HN project agent-memory-wiki allows temporarily online AI to choose what to leave behind and what not to. Comment: Worth watching if you want to know what AI thinks while working; authenticity of content is another story.

9. Letting AI Play "Werewolf". Developers had multiple AI models participate in an elimination-style game theory contest, observing differences in their alliance-building, deception, and voting performances. Comment: Game-theory evaluations are closer to the true nature of agents than Q&A; watching is fine, but don't treat it as a report card.

10. Calculating Costs: AI Agent "Thinking Watt" vs. Human Brain 20 Watts. An engineering blog compares AI inference energy consumption with human brain power usage; the human brain's 20-watt efficiency still makes the latest hardware envious. Comment: Another way to calculate compute anxiety; interesting to see, but estimates remain estimates.

Everyone Is Watching

Multiple AI coding agents working simultaneously stepping on each other's toes; open-source protocol Foremerge wants them to agree on signals before acting (GitHub / HN) · Engineers share that declaring "I am an expert" first changes output quality when working with AI (shimin.io / HN) · Design blog warns that risk-taking tendency in agents is a factory attribute, not learned post-factum (HAIPA / HN)

Tomorrow's Focus

① Anthropic's newly released Fable / Mythos 5.1 (cheaper, fewer restrictions) launched on Sep 1 has been live for one day. Wait for the first batch of user tests to see if "cheap" and "fewer restrictions" can stand together.

② Arena snapshot paused updates for one day on Sep 1. Watch two new faces in the next update: Alibaba Qwen 3.8 Max (#3 in code leaderboard, only 3,219 battles) and Tencent Hy4 preview (#5, 1,276 battles). See if rankings stabilize once samples accumulate.

③ Huawei openJiuwen's self-reported scores on the two exams: watch if third parties reproduce them. Only when reproduced successfully does it count as established.


Issue No. 035 · Preview Edition. Go to the official website for complete leaderboards and comments.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts