ReviewRadar · Global AI Evaluation Radar · 2026-08-18
🧪 Frontier Tech Review · Global AI Evaluation Radar
Tuesday, August 18, 2026 · Issue #20
5 minutes a day to see which AI software and hardware are worth your time
⭐ Key Updates
① AI is now better at "persuading people" than human experts
An arXiv study claims that AI systems have outperformed human experts in persuasion tasks. AI has no emotions, doesn't get tired, and can tailor rhetoric for each individual—persuasion is the underlying capability common to all business, politics, and scams. (Source: arXiv / Hacker News)
② Official deep dive into Claude's invisible watermarking
Anthropic explains the principle of invisible text watermarking: fine-tuning word selection probabilities embeds a "signature" into the text stream, which persists even after copying or rewriting. Quality impact and centralized verification rights are two points of controversy. (Source: TheVerge / Guardian)
③ Agent costs to rise 5x in five years
Enterprise-grade AI agent costs are expected to inflate by approximately 5 times by 2028. Engineering capabilities to save tokens will become key to scaling—this is also why the model routing layer is valued at $7 billion. (Source: The Register)
🏆 Leaderboard Flash
OpenRouter Latest Weekly Chart (8/10-8/16): Global usage reached 75.3 trillion tokens (+9.13%), with China accounting for 36.84 trillion, maintaining the global #1 spot for 16 consecutive weeks. DeepSeek V4 Flash official version topped the chart for two consecutive weeks with 11.2 trillion tokens, Tencent Hy3 came second with 9.95 trillion, and GPT-5.6 Luna entered the top three for the first time with 5.31 trillion. (Source: Daily Economic News 08-17)
Cost-Performance Chart: DeepSeek V4 Flash costs approximately 1/40th of GPT-5.6 Sol and 1/100th of Claude Fable 5 per task (Artificial Analysis data).
New Benchmark: CladBench launched—536 questions across 12 categories, specifically testing large models' mastery of UK and EU building regulations.
📋 Section Highlights
1. One-line config file makes AI code audit skip bugs — Prompt injection works on auditing Agents too; AI auditors also have "blind spot switches".
2. Vetted AI code is hard to justify — Code that runs and code that takes responsibility are two different things; half the bottleneck in implementing AI programming lies in "who signs off when things go wrong".
3. AI-generated Copilot Autofix exploited to breach Snowflake Jira — Auto-fix is a double-edged sword; AI-generated code must undergo manual review.
4. Do Not Trust, Continuously Verify — Default distrust of Agent outputs and step-by-step verification are prerequisites for scaled deployment.
5. Control charts make AI Agents cheaper — Monitoring Agents like a production line helps detect degradation early and reduce invalid calls.
6. Is reviewing AI even possible? — With weekly model updates and contaminated evaluation sets, evaluating AI requires a more dynamic anti-contamination system.
7. Training AI scientists to reproduce research — Having models read papers, run experiments, and verify reproducibility expands scientific research from generative to verifiable.
8. Local scanning of hidden Unicode in AI text — Zero-width characters and homoglyphs are easily caught; a self-cleaning device for the content ecosystem.
9. Doberman: AI watchdog preventing Claude from deleting databases — Pairing AI with an AI security guard; least privilege + real-time interception.
10. Post-mortem on OpenAI/Hugging Face hacks — Chaotic key management and overly broad supply chain permissions show basic security hygiene lagging behind model capabilities.
🔭 Trending Topics
- Anthropic becomes the "Apple of AI": highest revenue, but also the most expensive prices
- Canva's internal valuation cut by $10 billion
- GDS raises full-year guidance, expanding data centers by $10 billion this year
- Google Pixel new feature: AI analyzes battery usage
- MediaTek Day-0 adaptation for Alibaba Qwen 3.8
- ByteDance signs AI copyright agreement with Motion Picture Association
- Largest US indigenous tribe bans new data center construction
🗓️ Tomorrow's Watchlist
① See how DeepSeek's peak/off-peak pricing affects next week's OpenRouter weekly chart
② World Robot Conference opens on 8/19; Unitree's "Superman" robot makes its live debut
③ Claude's invisible watermark rolls out in gray scale; can third parties detect it?
④ After GPT-5.6 Luna enters the top three, the tug-of-war between open-source and closed-source call volumes will become clear
Views belong to the original authors; data is subject to official disclosures. This column focuses on the true level of AI software and hardware and does not constitute any investment advice.
Frontier Tech Review · Global AI Evaluation Radar | Shenzhen Frontier Tech Co., Ltd.
Physix Frontier