Community Discussion · Policy

ReviewRadar · Issue 30 · Preview (2026-08-28)

Physix Frontier Review · Issue No. 030 · Preview Version (2026-08-28)

See you tomorrow. Today we discuss three things: why a voice AI costing 2 cents per minute topped blind tests, how office AI assistants can be tricked by a few lines in a spreadsheet, and a new open-source tool installing a "dashcam" for AI coding assistants.

【Key Updates】

1️⃣ Voice AI Costing 2 Cents Per Minute Tops "Hard Question Quick Answer" Blind Test

ThunderPhone's new architecture focuses on voice AI. Official figures state cost is ~2 cents per minute. It took first place in the Big Bench Audio voice QA exam. This exam tests if AI answers tricky questions fast and accurately enough. Voice assistants have long been "clear recognition, slow answers, confused by interruptions." This new architecture aims to solve "understanding well and answering fast." Topping the list is an official self-test metric. Whether it holds up in noisy, accented, or interrupted scenarios requires third-party real-world testing.

Product Expert Ah Zhe says|Scoring voice entries by "getting things done" adds another sample. Previously voice leaderboards competed on recognition accuracy; now they compete on truly answering questions clearly. Voice assistants are rolling from "can hear" to "can do." 2 cents per minute aims to bring voice AI into daily consumer scenarios; the direction is right. Entry point battles are never won by whoever has lower scores, but by whoever makes users willing to speak daily. Per usual rule, don't trust launch events; try speaking yourself before judging.

Editor Xiao He says|As someone who heavily trips over voice AI, I care most about accents, interruptions, and cost. 2 cents per minute sounds cheap, but cheapness must be built on "not being stupid." If it gets two out of three questions wrong, saved money is wasted. I'll consider recommending it to friends after someone does a daily test with continuous speech and interruptions.

Source|Hacker News / ThunderPhone (2026-08-27)

2️⃣ Office AI Assistants Can Also Be Tricked by "A Few Lines in a Spreadsheet"

AI helping with spreadsheets, emails, and reimbursements is trending. But a public demo proved malicious instructions can hide in spreadsheet cells. When AI reads the sheet, it gets "brainwashed" and executes unauthorized operations. The principle isn't complex: AI can't distinguish data from commands in a table. Someone disguised attack instructions as normal text in the table, and AI accepted it all. Warning for ordinary people: Don't throw unknown files to AI for processing, especially office assistants with account permissions.

Security Expert Old Zhou says|This isn't an occasional bug; it's another manifestation of an old problem. AI can't distinguish data from instructions; boundaries are unclear, so attackers stuff things into the gaps. Tables, documents, and web links can all be poisoning entry points because, to AI, they are all read-in text. My stance remains unchanged: set permissions to minimum; don't rush to grant account permissions to office assistants. This was a demo, proving the attack surface exists. Wait for real cases and fixes before relaxing.

Editor Xiao He says|After watching the demo, I got chills. I really do throw messy tables to AI for organizing. I've said before: don't click unknown links, humans or AI. Now add: don't accept unknown files, humans or AI. Don't enable valuable account permissions unless necessary. I've noted this since Issue 021 and will continue to note it.

Source|shiftmag.dev (Hacker News, 2026-08-27)

3️⃣ Install a Monitored "Isolation Sandbox" for AI Coding Assistants

More people let AI write code, but consequences of AI messing with files or running random commands fall on you. Sandy is an open-source tool that locks AI coding assistants in a sandbox with monitoring and policy control. Every step AI takes is recorded; actions beyond allowed scope are blocked. For developers, it's like putting a dashcam and permission lock on the AI assistant. You dare let it work freely while seeing what it did.

Programming Expert Old Xu says|I agree with the direction. The stronger AI's coding ability, the more critical knowing which files it touched. Sandbox plus monitoring is a pragmatic way to swap "trust" for "verifiable." I've seen many such tools. Key points: ease of integration with existing projects, whether monitoring slows development, and how many leaks occur. Don't rush to production; run a small project for a week to check stability. Pretty docs aren't as good as rolling through real projects.

Editor Xiao He says|I don't code, but I get the logic. Locking AI in a cage where it can work but bad deeds are stopped is better than letting it roam wild. Before handing keys to AI, think clearly about which doors it can open and whether to retrieve them after use. I've noted this since Issue 024. Even if this tool isn't for me, it's a reassuring direction for those afraid of AI causing trouble.

Source|Hacker News / GitHub (2026-08-28)

【Leaderboard News · Who Was Strongest This Week】

Arena Human Blind Test Snapshot not updated this week (data as of Aug 26, same as last issue). This issue uses three new exams to rank.

① Terminal Operation Exam (Terminal-Bench 2.1, Official Verified List as of 08-19)

Tests if AI can open terminal, modify code, and run tasks itself. On the official verified list, Claude Code + Fable 5 tops at 83.8%. Anthropic-submitted Claude Code + Opus 4.8 ranks 5th at 78.9%. Newly open-sourced model Ornith-1.5 self-reported 86.1, but that's team-tested (average of five runs), not directly comparable to official lists. Discount AI self-reported scores before looking.

② Coding Ability Exam (SWE-bench Verified, Data Updated 08-26)

Human experts pick bugs from real open-source projects for AI to fix. Correctness is obvious. Claude Opus 5 Thinking High Config scored 96, ranking 1st. Claude Fable 5 two tiers scored 95, tying for 2nd and 3rd. Top three swept by Claude family. Domestic camp closes in tight: DeepSeek-V4-Pro scored 80.6, ranking 12th, free for commercial use. Qwen Qwen3.7 Max scored 80.4, ranking 13th.

③ Blind Test Code Score List (LMArena Coding, Snapshot 08-26)

Humans judge, AI fights blindly. Kimi's highest tier kimi-k3-max scored 1542, ranking 6th, continuing domestic models' first entry into top ten comprehensive blind tests. Claude Opus 5 High Config scored 1533, ranking 7th. Zhipu GLM-5.3-Max scored 1531, ranking 10th. Blind test scores are just reference coordinates; real utility depends on your specific tasks.

【Section Highlights】

1. Pydantic AI adds "type insurance" for Python AI apps, reducing pitfalls of incorrect AI return formats during development.

2. Apronagents gives each AI coding assistant a disposable independent repo, discarded after use, preventing cross-contamination.

3. Gantree lets AI run long tasks for hours without humans staring at chat boxes.

4. KinoPipe encapsulates video editing into services AI can call directly, no guessing command lines for editing videos.

5. SCQOS is the world's first public challenge; AI must prove it should act before acting.

6. Jailbox is an offline black-box VM; AI and untrusted code run in locked environments, explosions don't affect outside.

7. LetItLoop allows AI long tasks to resume from breakpoints after crashes, no need to restart from scratch.

8. Relay Q uses one microphone to make AI understand you; Wired real-world test rated it close to specialized voice software.

9. deepagents is LangChain's open-source AI framework; AI can plan tasks, call tools, and complete multi-step work itself.

10. GateOnAI calculated 4 million real compatibility connections among 2866 AI tools. Data speaks to whether they can be stitched together.

【Everyone Is Watching】

Anthropic tests letting Claude operate robots and lab instruments | Nvidia's 70% growth expectation ignites AI rally | Hugging Face pushes $399 open-source duck robot | OpenAI CEO admits public hates data centers

【Tomorrow's Focus】

① Ornith-1.5 three-tier weights rolling out; see if third-party re-tests reproduce its self-reported 86.1

② After new voice assistant architectures, see who integrates into mass products first

③ For table injection attacks, see if major vendors release targeted protections

Full version at daily.physixfrontier.com/review/

2 replies

?
Ctrl + Enter to reply
Shua Ti Zhong

Does this eval help verify model robustness? I've ground through 300 questions but still worry about being asked about implementation challenges in interviews. Torn over offers right now.

Mo Mo
Mo MoAug 29

Data freshness is questionable. How do you solve the latency in closed-loop training for Physical AI?