Physix Frontier Evaluation · 2026-09-13
Issue #046 · Preview Edition. Others do reviews; we build the radar for reviews. Today's Arena blind test leaderboard didn't refresh, so the quick update switched to Artificial Analysis' real-time snapshot. All numbers were grabbed fresh this morning.
Key Updates
1. Informatics Olympiad Papers Become the Watershed for AI Capability
Knowledge-based leaderboards can no longer stump flagship models. Vals.ai turned past International Olympiad in Informatics (IOI) problems into a test paper. GPT-6 Astra solved all three years' worth of problems. The top tier has noticeably fewer members than on knowledge leaderboards, with a cliff-like drop in scores for the middle tier.
Question setting and grading are entirely handled by Vals.ai's internal process. Self-reported scores should be discounted initially; wait for independent re-runs before taking them seriously.
Programming Expert · Old Xu says | Olympiad problems test the ability to break down unseen questions into algorithms, which is much harder to fake than memorizing question banks. I approve of the direction. But per the old rules, wait until two or three independent parties re-run the same set of problems and the scores match, then use it as a basis for selection.
Editor Xiao He says | Previously, exams were like open-book tests where everyone could copy their way to a 95. Now this paper is like math olympiads. In the future, when you see "Model X tops Leaderboard Y," first ask what was tested. A leaderboard where everyone gets high scores proves nothing.
Source | Vals.ai (2026-09-12)
2. Your Skills for Evaluating AI Employees Might Not Be the Skills You Actually Hire
Anuclei published an article that sparked heated discussion in programmer circles. Teams spent months building evaluations for agents and getting high scores, but after going live, they ran a completely different configuration: different prompt versions, smaller model variants, and added/removed permissions. The agent during evaluation is fundamentally not the same thing as the agent doing the work.
Security Expert · Old Zhou says | This is a structural problem; individual companies cutting corners is just a symptom. Evaluation environments default to being lenient, while production environments have complex permissions, and no one is responsible for the drift between the two. When picking agent products, you can directly ask, "Does the scorecard you publish run on the exact version I'm holding?"
Editor Xiao He says | Translated into plain language: they show off the top-spec model during the interview, but hand you the budget cut-down version for the job. Last issue, I said "For fully automated AI operations, first ask if they dare to publish audit logs." This article adds the second half: you also need to ask whether the published logs are from the interview or from the actual job.
Source | Anuclei (2026-09-13)
3. $4,000 Chinese Robot Dog: Journalist Ran It for 42 Days First
A senior Ars Technica journalist bought a Chinese-made quadruped robot dog out of pocket and wrote a long-term review after using it for over a month. It can follow you on walks, carry cameras, and dodge people in the yard on its own, but there were plenty of crash sites too: it got stuck staring at a vacuum cleaner for half an hour and simply quit working in the rain.
Wearables Expert · Akai says | Continuing my judgment from August 17 regarding the $16,000 lawn-mowing robot, the real barrier for smart hardware lies in installation, power supply, waterproofing, and other parts not covered in the manual; performance is just the entry ticket. What moved me most about this piece was the journalist's description of the psychological shift on "Day 30 of gathering dust." That is the data consumers need most.
Editor Xiao He says | Before buying, confirm if it can return to its dock to charge automatically. Things you have to manually carry back every day will sit idle within a week. This unit's app is entirely in English, and firmware updates feel like opening a mystery box. From a domestic user's perspective, it's a different story since local versions come with Chinese after-sales support. Waiting for a domestic long-term review before deciding whether to save up.
Source | Ars Technica (2026-09-12)
Leaderboard Quick Update · Who Was Strongest This Morning (Artificial Analysis Real-Time Snapshot)
The top tier of the Intelligence Index is as follows.
| # | Model | Vendor | Total Score |
|---|---|---|---|
| 1 | Claude Fable 5.1 (Max w/ fallback) | Anthropic | 53 |
| 2 | GPT-6 Astra (Max) | OpenAI | 53 |
| 3 | Claude Opus 5 (Max) | Anthropic | 51 |
The first tier is crowded around 53 points. The strongest configurations from both companies differ by only fractions, making it hard to tell who wins by eye. This leaderboard condenses dozens of exams into a single total score; it's good enough for a rough ranking, but choosing tools still requires checking specific metrics based on your scenario.
In the open-weight leaderboard (models you can download and run locally), GLM-5.3 (Max) leads with 45, followed by Kimi K3 (Max) at 44, and GLM-5.3-Flash at 42. The top three spots are swept by two Chinese companies, trailing closed-source flagships by only about 8 points. A year ago, that gap was over 15 points.
The speed champion is Celeris-1 at 1313.8 words per second. The cost-efficiency champion is Granite 4.2 3B at one cent per task. Fast, cheap, and strong are three different leaderboards.
Data taken from artificialanalysis.ai/leaderboards/models. The latest snapshot of the Arena Human Blind Test leaderboard remains September 12, same as last issue, so it wasn't used today.
Section Highlights (10 Items)
- In the era of AI writing code, "fix later" TODO comments have become time bombs. Agents execute TODOs as instructions, and "temporary key to delete later" actually makes it into production. After working with AI for too long, humans forget that it literally executes every sentence.
- Metrum AI Router goes open source, learning "which model handles which task": simple steps go to cheap models, hard problems stay with expensive ones. Saving money relies on changing scheduling, not changing brains.
- "Father of Claude Code" Cherny states that developers' core responsibility has changed: it's no longer typing every line yourself, but guarding code quality. When the tool-maker stamps approval, the wind direction is set.
- Engineer Sean Goedecke plays devil's advocate: "Stop building tools solely for AI agents." Agents are transient forms; stripping human interfaces for their sake puts the cart before the horse. Hundreds of arguments on HN.
- The Wall Street Journal reviews AI tackling millennium math problems. A brain capable of solving thousand-year-old puzzles can easily find system vulnerabilities. Exams need to switch from humanities to sciences.
- Wired's final security weekly summarizes Claude abuse archives, ranging from intrusion scripts to bioweapons. Daring to expose missed records is more credible than claiming safety.
- A single-person longitudinal experiment shows that AI relying solely on static file storage cannot sustain long-term collaboration; dynamic state (where things stand) is the missing piece. Sample size is only 1, but note the direction.
- Show HN: A free AI game production handbook. Cover principle: No Fabricated Numbers.
- Wired interviews mathematician Strogatz, who tears up discussing AI progress. He also confirms that AI-generated "proofs" still require line-by-line human verification. Shock is shock; verification is verification.
- Ant Group's GPASS upgrades to "Lingying," installing an agent-native OS on AI glasses. Future tests of AI wearables will first look at whose foundational layer they run on.
Everyone Is Watching
Vals.ai IOI leaderboard launch, agent "eval version ≠ deployment version" going viral, 42-day robot dog long-term test, AA real-time leaderboard tie at 53 points, Wired abuse archive finale.
Tomorrow's Focus
- If the Arena blind test leaderboard refreshes with a September 13 snapshot, compare it against this issue's AA real-time leaderboard to see who takes the top spot.
- First third-party re-runs of the Vals.ai IOI leaderboard. Key focus: Can GPT-6 Astra's "full solution" be independently reproduced?
- First batch of partner terminals for the Lingying open platform. Confirm which AI glasses get access to real-device evaluation slots.
Physix Frontier