ReviewRadar · 2026-09-10
Issue #043 · Preview Edition. Others do reviews; we build the radar for reviews.
Key Updates
1. Why is AI work so expensive? Someone installed a "burn rate meter" on it
Nowadays, many AIs can flip through files, look up info, and write code for you. Costs are usually calculated by "how much text you feed it." This new tool did some real-world math: when an AI flips through entire books to find answers, the actual amount of text it swallows is 16 to 47 times what the answer actually needs—equivalent to buying a whole book for one page of content.
Lao Xu (programming expert) said: I previously stated, "When AI output increases, verification must keep up," but today I need to add: when output increases, wasted reading volume must also be managed. When humans look for one page in a book, they go straight to the table of contents; AI uses the brute-force method of reading the whole thing. However, it just appeared today and is still a rough tool—wait for third-party tests before integrating it into serious projects.
Xiao He said: I don't understand the billing formula, but I get this "meter": for the same AI, some tasks cost pennies while others cost several dollars. It's not always because the problem is hard; sometimes it's stupidly re-reading the book repeatedly. Let's wait for the first batch of real users to share their bills. Remember one thing: making the AI read less useless material saves money.
2. Want to install plugins for AI but afraid of getting hacked? Pass this "security checkpoint" first
AI can now connect to various third-party plugins, like installing apps on a phone, with no one vetting quality or safety. This open-source tool puts suspicious plugins in an isolated sandbox first: checking if they steal your data or have known risks, then summarizing results on one page for you.
Lao Zhou (security expert) said: When reviewing similar tools before, I said that interception based on "environment control" is how least privilege is truly implemented. Tripwire follows this path: it doesn't reason with plugins; it locks them in a sandbox to see what they do. But detection rules are written by humans, and coverage determines the ceiling; this is a personally maintained project. As usual: don't rush to grant permissions; the security checkpoint itself needs to pass inspection first.
Xiao He said: I still remember the Copilot vulnerability where "clicking confirm could be bypassed by links," so I respect anything that says "scan before connecting plugins." But words like "isolated environment" give me a headache immediately. If I can't use it in three steps, I'll probably give up.
3. Afraid of picking the wrong "exam paper" for voice AI? Someone compiled 36 benchmarks into a directory
Whether voice AI (assistants that answer calls or help book locations) is good depends on many test items—accuracy, turn-taking logic, task completion, etc., which were previously scattered everywhere. This website organizes 36 public benchmarks into a directory categorized by 6 capabilities: if you want to find tests for "judging when to speak," click in to see who created it, who took it, and the scores. For companies selecting voice solutions, this is like having an overview map.
A Zhe (voice product expert) said: I've always said the watershed for voice AI lies in latency, interruption, and barge-in. Flipping through this directory, I first counted the proportion of "turn-taking" benchmarks—if the exam focus shifts to "who should speak," it means the industry finally grasped the key issue. But as usual: the directory doesn't generate scores itself. You must click in to check each exam's methodology, whether vendors gamed the questions, or if it's self-promotion. Before selection, run candidates against your own customer service recordings; that's harder than any leaderboard.
Xiao He said: I don't understand tech, but the scenario of "interrupting a customer service bot" instantly raises my blood pressure—this directory is quite comprehensive, saved it. I hope the next step allows one-click testing with your own recordings; otherwise, no matter how complete, it's just a directory.
Leaderboard Flash (Snapshot stopped on Sept 9, not yet updated)
- Text Blind Test Leaderboard: Anthropic occupies six of the top ten spots, nearly monopolizing it. Scores come from anonymous one-on-one human votes on "who answered better." Ranks 3 and 5 have fewer than 5,000 votes, so positions are volatile.
- AI Worker Leaderboard (real tasks, completion rate): Claude Fable 5.1 Max is first. Tencent Hunyuan Hy4 Preview ranks in the top ten, highest among domestic models. Note that "completion rates" are generally around 20%—it's too early to fully let AI handle tasks independently.
- Coding Leaderboard: OpenAI's new model temporarily leads but has only accumulated 1,810 votes; the 32-point gap with #2 isn't settled. Alibaba Qwen and Kimi squeezed into the top five; domestic coding capability is really rising.
Section Highlights (One sentence each)
- Scholars cross-checked "AI-written review comments" item by item: AI is good at spotting writing flaws but often misses scientific holes—final human oversight cannot be skipped.
- Nature had AI "rediscover" relativity: it can recite textbooks, but reasoning from scratch often gets stuck halfway—solving problems and creating problems are two different things.
- GitHub may be discussing adding "AI-generated" labels to content, allowing admins to bulk-filter spam—AI content is starting to require identity disclosure.
- A project is building a small tool: after AI breaks code, verify old backups are intact and save current state before rolling back—like saving twice before shutting down.
- Someone turned AI instructions and materials into reusable "packets," avoiding re-explaining everything for each project—the pain point is real, but such small tools risk being crushed by big players anytime; wait and see.
- A full-time mom used "teaching kids not to bite" to explain "AI alignment," stunning engineers—good analogies spread faster than papers.
- A veteran developer with 25 years experience posted "PHP is dead, blame AI": languages frequently seen by AI become more popular; the tone sounds like arguing, but the trend seems real.
- An open-source coding assistant running entirely locally launched: code never leaves your computer—a real need for privacy-conscious developers—just launched today, bookmark it and wait for honest feedback.
- Foreign media compiled a complete guide to planning travel with AI: let AI fetch info, keep decision-making power yourself—this usage works for any assistant.
- Someone built a "comment padding detector": AI-added code comments are mostly "correct but useless," and this tool specifically separates them from human-written comments—the most expensive part of code review isn't errors, but humans reading AI's fluff.
Trending
Anthropic researcher resigns warning AI race is "betting humanity's lives" | Apple releases privacy statement explaining Watch "Sound Recognition" | Suno v6 switches to licensed music training to address copyright lawsuits | Tencent Hy4 free trial expires tonight | iPhone 18 Pro photo anti-forgery labeling launches.
Tomorrow's Focus
1. iOS 27 confirmed to push officially next Monday (Sept 14), with new Siri launching alongside: Don't jump on day one; wait for first-batch user tests.
2. Hunyuan Hy4 Preview free trial expires tonight at 23:59; new user benefits adjust tomorrow: If you want to grab meeting minutes deals, export results tonight.
3. Arena leaderboard snapshots stopped on Sept 9: Watch for new data tomorrow, especially if the #1 and #2 coding models maintain their lead once sample sizes increase.
Physix Frontier