ReviewRadar · 2026-09-15
Issue #048 · Preview Edition. While others do reviews, we build the radar for reviews—today’s issue tracks three things: Edge AI gets its first public report card, "Third-party security evaluation" becomes an industry consensus topic, and robot vacuums enter the long-term testing era.
Key Updates
1. Want to run AI on your old computer? Check out this "Bare-metal AI" report card first
Someone just set up a "Bare-metal AI" leaderboard on GitHub: it specifically catalogs ultra-lightweight AI engines written by hand in C/C++/Rust without relying on big frameworks, marketed as zero-dependency with sub-millisecond response times—the batch of "little powerhouses" that let AI run directly on your devices without GPU server farms have been ranked under the same test paper for the first time, complete with speed benchmarks and supported model indexes.
Lao Xu (Programming): The questions and scoring are all defined by the repo owner; it smells heavily of marketing. But edge AI definitely lacks a public report card. Don't rush into selection; pick the top two from the list and run them against your own workloads for a week.
Xiao He: Beginners only care if installation is a hassle or if they need to buy new hardware. Save it for now, don't touch it yet; wait until someone tests how sluggish it runs on a regular laptop.
2. Who decides if AI is safe? Why should manufacturers grade their own homework?
Bloomberg's show yesterday specifically discussed whether independent testing can make AI safer. The background is the heated "slow down the frontier" debate over the past few days: Anthropic's CEO called for hitting the brakes, OpenAI and Musk followed suit, and eventually everyone pointed in the same direction—rather than trusting self-reported scorecards from vendors, better to leave it to external agencies.
Lao Zhou (Security): Independent testing is the right direction, but watch three things—are the test questions public, is the scoring reproducible, and is it testing the version you actually have? If they test a special edition, the credibility is no different from self-reported scores.
Xiao He: Previously athletes awarded themselves medals; now there's a proposal to hire outside referees. Wait until the "safety checkup reports" hold steady rankings for two consecutive issues and raw data is downloadable before using them as a basis for choosing tools.
3. Robot vacuums also enter the "exam era": Wired long-term tests Roborock Qrevo 2, declaring the budget king has changed hands?
Wired published a long-term review of the new Roborock Qrevo 2, with the title directly asking "New budget robot vacuum king?" Specs are easy to write, but long-term tests reveal the truth: obstacle avoidance, corner cleaning, and whether the base station's mop washing smells bad—all are results editors got from actual home use.
A Kai (Hardware): The primary reason for collecting dust is whether installation and charging are smooth, not suction parameters. Future battles for the throne must be settled by public performance data—the domestic-led 11-item home robot testing standards have just been established, so praise from a single media outlet only counts as half the battle.
Xiao He: Note down "budget king" for now; wait for long-use posts spanning more than three months. That's the original buyer's showcase.
Leaderboard Flash
Comprehensive Intelligence Index (Artificial Analysis pulled live this morning): Claude Fable 5.1 (both modes) and GPT-6 Astra tie at the top with 53 points across three entries, followed closely by Claude Opus 5 at 51 points. The open-source camp switches flags—Zhipu GLM-5.3 (45 points) surpasses Moonshot AI Kimi K3 (44 points) for the first time. Fastest is Celeris-1, spitting out 1418 tokens per second.
Chinese Authority List (SuperCLUE, updated Sept 11): DeepSeek V4.1-Flash, released last week, entered the list immediately and jumped to second place among domestic models (71.81 points); GLM-5.3 surged nearly 8 points from 63.27 to jump to 3rd. On LMArena's global company leaderboard, Zhipu rose from 6th to 5th driven by GLM-5.3, pushing Moonshot AI down.
AI That Gets Work Done (Data as of Sept 8-10): In the OSWorld exam where "AI operates the computer for you," Chinese team Shizai Agent topped the list with 90.2% (second place Claude Fable 5 scored 86.0%); on the doctor's exam HealthBench, Claude Fable 5 scored 0.660 vs. human doctor baseline of 0.437—the test is too easy, don't think AI can diagnose diseases; BenchAlign comprehensive top three remain Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5.
Note: The Arena human blind test snapshot is still the Sept 14 batch, same as last issue; the leaderboard hasn't updated yet, so no new changes are cited in this issue.
Section Highlights
1. Break down AI bills into four segments—"loading, thinking, scheduling, outputting"—to calculate costs. If you can't distinguish effective thinking from AI flipping through the whole book, you won't save money.
2. arXiv paper: AI helps organize interview texts from tens of thousands of people; social sciences start competing fiercely, also discussing how to prevent AI from smuggling biases during induction.
3. Build software without coding skills: Write requirements and acceptance criteria into fixed workflows, let AI produce according to the process, humans only check at nodes—more reliable than chatting and leaving it to luck.
4. Security researcher demonstrates hijacking customer service AI with one email: Hide instructions in the email body, and AI can't distinguish between "customer letter" and "system command." This boundary-blurring issue is chronic, not occasional.
5. Too good a memory is also a disease for AI assistants: Long-term memory treats expired conclusions as truth and continues using them with confident tone. Set expiration dates for AI memory.
6. Security firm warning: AI mass-produces "boss urging payment emails," eliminating the old flaw of stiff grammar. Always call back the person via phone to verify any transfer requests.
7. Researchers publish a list of 1,325 GitHub codebases with traces of AI assistance—for the first time, there's a large-sample database for studying the survival rate of AI-written code.
8. Show HN new toy Otis: Install and it calls local models to work, no cloud connection, no registration; "data never leaves the machine" is its entire selling point.
9. Threshyr 1.2.1: Time tracker usable offline, edge AI categorizes your day into "time blocks." For always-on background tools like this, try it on a backup device for two weeks first.
10. ProGantt: Connects Gantt charts to AI (MCP interface), allowing AI to view schedules and modify tasks. But AI config changes don't leave error logs; always check the change list before merging.
Everyone's Watching
Zhipu rises to 5th in company leaderboard, pushing Moonshot AI down | Anthropic occupies six of the top eight spots in Arena text leaderboard (snapshot not updated) | DeepSeek V4.1-Flash enters SuperCLUE and immediately ranks 2nd among domestic models | Open-source gap narrows to within 3 points, refresh could swap the leader anytime
Tomorrow's Focus
① Next updates for SuperCLUE and LMArena: Will GLM-5.3 hold onto 3rd place? Note new rankings but don't commit yet.
② Third-party re-runs of DeepSeek V4.1-Flash: Official claims "rebuilt from scratch, faster and more accurate"; wait for independent exams to re-test before believing it.
③ If Arena snapshot updates, watch claude-fable-5.1-max (only 5,783 battles, thinnest sample)—will votes accumulate enough for confidence?
④ Three-way fight in open-source leaderboard: GLM-5.3, Kimi K3, and GLM-5.3-Flash are within 3 points; any refresh could swap the leader.
——Full leaderboard data and dual-expert commentary available on today's Review Issue page.
Physix Frontier