ReviewRadar · 2026-08-30
Physical World Frontier Review · 2026-08-30
(Issue #032 · Preview Edition)
Others do reviews; we do radar for reviews. This issue focuses on three things: Google's new model climbing three leaderboards in two days, someone creating a criminal record database for AI hallucinations, and the craft of "putting AI inside your own computer."
1. Google Gemini 3.7 Flash Climbs Three Leaderboards in Two Days: Chat #9, Coding #8, Task Execution #20
What it can help you do: This is Google's volume-oriented AI released on August 13, focusing on cheap and fast. New data published by Arena's official account on August 28: Chat blind test score 1490, ranked #9 (leaderboard snapshot as of 08-26, 5,718 battle samples); WebDev coding leaderboard jumped from #19 to #8; Agent leaderboard (testing if AI can run tasks and get things done) entered at #20, with "net improvement" +3.4 and "task completion rate" +10%, ranking #5 overall. Compared to predecessor Gemini 3.6 Flash: Agent #35, Coding #21—a leap with generational change. Pricing is aggressive too: Reading 1 million words costs ~$0.75, writing ~$3.75 (AI billed per token, one token ≈ 0.75 Chinese characters), limited-time price. The August LLM weekly report also noted it outputs 343.7 tokens per second, fastest overall, ranking #10 in comprehensive scores (compiled by third party, cited from Zhihu).
Expert Commentary · Lao Xu (Programming): Last issue I warned: wait for rankings to stabilize for a round or two before judging. Today's report still has thin samples—Chat leaderboard has 5,718 battles, while veterans often have tens of thousands; Coding leaderboard has only 2,552 for it. The direction is truly strong: predecessor Gemini 3.6 Flash ranked #21 in coding, yet one generation lifted it into the top ten, showing Google got the "fast and cheap" line right. But old rule: treat vendor launch stats and newly entered ranks as unverified until tested. Run real tasks for a week before integrating into projects to see where it breaks.
Editor Xiao He: Price is indeed attractive, but I continue the approach from Issue #031: watch for volatility in newly ranked models. Also, "speed" is most intuitive for ordinary users like me—no waiting for spinning wheels when replying. I'll try writing weekly reports with it first, keeping important tasks for the provider I've used for six months.
2. Someone Created a "Criminal Record" Database for AI Hallucinations: Latest Entry from Aug 27 US Court Records
What it can help you do: A website called the AI Hallucination Case Library collects real records of AI "confidently making things up": fabricated legal precedents, non-existent citations, invented facts, recording parties, times, and sources item by item. On August 29, it surfaced on Hacker News; the latest entry is from an August 27 lawsuit in Florida, USA, where someone used AI-fabricated material as evidence. The value of such libraries lies in aggregating scattered failure cases into a searchable ledger—next time you see AI confidently listing "sources," check against this ledger first.
Expert Commentary · A Zhe (Product): I position these as "buyer review archives." Model quality is largely defined by vendors; but cases exposed in court or debunked by journalists cannot be PR'd away. Reclaiming scoring power from vendors relies not on another benchmark list, but on negative lists with case numbers. It doesn't tell you which model is strongest; it tells you all models may bury landmines in your most serious report.
Editor Xiao He: In Issue #028 I said "check buyer reviews before buying"; now even failure showcases are archived. When using AI to write, verify every material, link, and number provided before forwarding—my old habit: verify before trusting AI.
3. Is Putting AI "Inside Your Own Computer" Hard? Wired Published a Practical Guide from Scratch
What it can help you do: Wired published a tutorial on August 29 teaching how to run a chat AI offline on your own computer: selecting models, installing software, configuring hardware, step-by-step. Offline usability and local privacy are the biggest selling points; the cost is that small models runnable on home computers have visibly inferior comprehension and knowledge compared to cloud flagships.
Expert Commentary · A Kai (Wearables): I've tracked local AI since Meta open-sourced the 30B model on August 11, 2026. My judgment hasn't changed: consumer GPUs running open-source models is a watershed moment, greatly increasing playability, but the gap between small models and cloud flagships is real. This tutorial is worth saving for clarifying "what hardware achieves what level"—choose tiers based on your PC specs; don't challenge 70B models with ultrabooks. If it runs poorly, don't blame the AI.
Editor Xiao He: Offline capability is a genuine need, but I repeat my phrase from Issue #017: fear that local versions secretly swap in a "smaller brain" and become dumber. Try installing on a backup machine per the tutorial first; hold off on tinkering with the main machine.
Leaderboard Updates · Who's Strongest This Week
Chat Blind Test Leaderboard (Data as of 2026-08-26)
- 1. Claude Fable 5 (Anthropic) 1508
- 2. Claude Opus 4.6 High (Anthropic) 1504
- 3. Claude Opus 4.7 High (Anthropic) 1502
- 4. Muse Spark 1.2 xHigh (Meta) 1498
- 5. Claude Opus 4.6 (Anthropic) 1497
- 9. Gemini 3.7 Flash High (Google) 1490, newly entered top ten
In human-judged, anonymous blind tests, Anthropic holds six of the top ten seats, Meta two. Google's newcomer Gemini 3.7 Flash squeezed into #9 with 1490 points. Blind test scores involve humans voting for winners between two anonymous models, converted to scores. The top 8 differ by less than 20 points; mutual wins/losses are normal.
Agent Leaderboard · Testing If AI Gets Things Done (Data as of 2026-08-26)
- 1. Claude Opus 5 High (Anthropic) Net Improvement 12.73
- 2. Claude Opus 5 Max (Anthropic) 12.41
- 3. Claude Fable 5 High (Anthropic) 11.62
- 4. GPT-5.6 Sol xHigh (OpenAI) 10.31
- 6. Kimi K3 Max (Moonshot AI) 9.53
This board doesn't ask "who speaks beautifully," but "if tasked, did it complete the job?" The Claude family sweeps the top three. Domestic Kimi K3 ranks #6, with the highest "task completion rate" (18.24) but low "instruction adherence" (2.31)—it works well but doesn't always listen. Newly entered Gemini 3.7 Flash is temporarily #20, with few samples.
Coding Leaderboard (Data as of 2026-08-26)
- 1. Claude Opus 5 Max (Anthropic) 1691
- 2. Kimi K3 Max (Moonshot AI) 1674
- 3. Qwen3.8 Max (Alibaba) 1669
- 4. Claude Opus 5 High (Anthropic) 1663
- 10. Gemini 3.7 Flash High (Google) 1587, rose to #8 in WebDev sub-leaderboard on Aug 28
Two open-source contenders (Kimi K3, Qwen3.8 Max) broke into the top three on the coding board, pushing paid closed-source models behind them.
Section Highlights · 10 Review Trends Worth Your Time
1. Simurg: Free Search Tool for AI – Prefers Interruption Over Hallucination. Newly launched open-source tool providing free web search for AI agents, selling point: "abort hallucination." If no source supports the answer, it stops rather than continuing to fabricate. One-line comment: "Stop rather than lie" is a targeted approach; tool is new, interception effectiveness awaits third-party testing.
2. Documentation.ai.md: Proposes Standard for "Docs Written for AI." Open proposal: Software docs shouldn't just be for humans, but for users' AI assistants to operate directly—software now has two readers: those deciding usage, and agents doing the work. One-line comment: Docs are evolving a new species "read by machines"; whoever standardizes first owns the entry point.
3. Programmer's Long Post: Beyond Coding, First Real Time-Saving AI Use is Grocery Shopping. A developer reviewed that outside coding, the only stable time-saving AI use is planning weekly grocery lists—considering budget, dietary restrictions, and expiring fridge items. One-line comment: AI's practical value often hides in unglamorous chores; personal logs are more credible than press releases.
4. BoqCalc: AI Prices 500-Line Engineering Quotes, Focuses on "No Fabricated Unit Prices." Construction industry AI pipeline: upload bill of quantities, it calculates costs, flags hidden risks, locking quote unit prices in verifiable data, preventing model improvisation. One-line comment: Another design sample preferring "no calculation over wrong calculation"; engineering errors are costly, wait for real project cases before use.
5. AI Begins Touching Cryptography Foundations. Security podcast invited cryptographer Chris Peikert to discuss recent progress: AI beginning to automate proofs on mathematical problems like "lattices"—lattice hardness assumptions underpin post-quantum cryptography. AI can help verify, theoretically also find flaws. One-line comment: AI moving from "solving problems" to "proving theorems" shocks then delights crypto circles; "AI breaking encryption" remains a hypothetical drill, not reality.
6. Security Researcher Public Call: Send Me Your AI Agent's Weirdest Logs. Researcher soliciting "weirdest agent logs" from users—records of AI going off-track or acting autonomously while working, to study real behavioral boundaries. One-line comment: Agents' most honest capability material isn't demos, but failure process logs; folk samples are scarcer than papers.
7. Veteran E-book Manager Calibre Adds AI-Generated Covers. Select a book, let AI generate cover art based on title/content without design skills. One-line comment: AI image gen is becoming default feature in daily software; check copyright terms for public use.
8. "Ask AI" Button: Websites Starting to Embed AI Entries Like Social Shares. No-registration widget: Site owners generate "Ask AI What This Site Says" buttons; visitors click to chat with their own AI assistant. One-line comment: Users already habitually ask AI before clicking websites; sites are embedding AI...
Physix Frontier