ReviewRadar · 2026-08-29
Physix Frontier Reviews · 2026-08-29
5 minutes a day to understand which AI tools are worth using. Today is the preview version of Issue #031.
The human blind test leaderboard hasn't been updated for a few days; the data is still stuck on August 26. However, Google dropped a bombshell on the evening of August 28, bringing this issue's rankings fully to life.
Key Updates
1. Google AI jumps onto three leaderboards overnight
The official Arena (formerly LMArena) account announced that Gemini 3.7 Flash (High) has landed on three leaderboards. It ranks 20th overall in the Agentic Leaderboard; the previous generation, Gemini 3.6 Flash (High), was ranked 35th, so this is a jump of 15 positions. Its tier for long-horizon tasks is now on par with Claude Sonnet 4.6 and GPT-5.6 Terra (xHigh). In the WebDev sub-leaderboard for coding, it jumped from 19th to 8th. On the Text Chat leaderboard, it sits at 9th with a score of 1490. Regarding pricing, until the end of 2026, input costs $0.75 per million tokens and output costs $3.75. Starting January 1, 2027, prices will revert to $1.50 and $7.50 respectively.
Programming Expert · Old Xu says | Moving from 35th to 20th on the Agentic Leaderboard and 19th to 8th on WebDev represents real ranking changes. But these new ranks have only been on the board for a day or two, and the sample size isn't sufficient yet. Let's wait for it to stabilize for a round or two first. The vendor's own report cards from the batch released on August 13 (FrontierCode 43.6%, DeepSWE 65.3%) should be discounted as usual. If you want to verify, run your own real projects through it.
Editor Xiao He says | So Google's latest model improved its rankings across all three areas: working, coding, and chatting, and there's a price discount until year-end. However, I just learned last issue that new leaderboard rankings fluctuate, so don't rush to switch your entire stack to Google. Public leaderboards are essentially buyer reviews before purchase. The leaderboard data is current as of August 26, while the new rankings were announced on August 28, leaving a two-day gap.
2. When AI scans code for vulnerabilities, 44% of reported "critical flaws" are false alarms
US telecom giant Comcast tested AI vulnerability scanning on their own 170 million lines of code. Of the "high-risk critical vulnerabilities" reported by the AI, 44% were false positives. They also compared two modes: exploratory evaluation, where the AI operates freely, and deterministic evaluation, where the AI is given explicit instructions and context. The results between the two differed significantly.
Security Expert · Old Zhou says | A 44% false positive rate is worth remembering. AI scanners can't distinguish between "looks like a vulnerability" and "is actually a vulnerability." It's still that old problem: unclear judgment boundaries. But separating the two evaluation modes in practice shows that the clearer the instructions given to the AI, the fewer the false positives. The conclusion is simple: results scanned by AI must undergo manual review.
Editor Xiao He says | I thought whatever the AI reported was true, but nearly half of it was crying wolf. This aligns with what I learned in Issue #029: verify AI claims before trusting them. From now on, if someone hands you a report titled "AI Discovered Security Vulnerabilities," the first principle is: don't panic, find someone to confirm if it's a false positive.
3. Do AIs have political stances? Someone put various large models through a political compass test
A website called AI Political Compass dragged various large models through the original "Political Compass" test, marking each model's position on the political coordinate system. You can click on company names to compare them individually. The resulting positions vary for every single one.
Product Expert · Ah Zhe says | Models aren't neutral answer machines; their stance is factory-set, written into the training data. The value of this test is turning invisible biases into visible coordinates. Choosing a model shouldn't just rely on benchmark scores; you also need to see how it answers sensitive questions. That itself is part of the product.
Editor Xiao He says | When I ask AI questions, I usually assume its answer is the standard truth, never considering that the answers carry the bias of the training data behind them. Putting different companies side-by-side reveals quite obvious differences. In the future, when asking about social topics, I'll check multiple sources and not treat AI answers as stance-free standard truths.
Leaderboard Flash · Data as of 2026-08-26 (Official update pending)
- Text Chat Leaderboard: Claude Fable 5 (1508 points) leads. Anthropic holds four of the top five spots, with Meta's Muse Spark 1.2 sandwiched in between. In human blind tests, the Claude family almost monopolizes the board.
- Agentic Leaderboard: Claude Opus 5 (High) leads with a net improvement of 12.73 points. Anthropic holds four of the top five spots. OpenAI's GPT-5.6 Sol ranks 4th. Domestic Kimi K3 ranks 6th.
- Coding Leaderboard: Claude Opus 5 (Max) is first with 1691 points. Open-source contenders Kimi K3 and Qwen3.8 Max broke into the top three.
- New in this issue: Gemini 3.7 Flash landed at 8th on the WebDev sub-leaderboard on August 28. See Key Updates for details.
Section Highlights (10 items)
1. Did AI say the code is done? This tool helps you check if it's ready for production (AI Shipcheck). Runs local checks to see if AI-generated code is truly deployable. No registration required, no source code uploaded. Turns "AI said it's good" into "verifiable." The direction is right, but since the tool itself was written with AI assistance, try it on small projects first.
2. Wanting to steal sensitive data? This open-source gatekeeper stops AI agents (Weir). Reads the agent's own tracking logs. If sensitive data flows to unauthorized places, it fails the build. Turns "the environment decides" from a concept into a tool.
3. Pass a safety evaluation before sending commands to factory robots (URML). Uses a human-readable mini-language to describe robot intent, then determines if execution is allowed on real devices. For physical AI to go live, we first need "auditable intent." The direction is very positive, but implementation is still early.
4. Open-source voice foundation, giving AI a mouth that can interrupt and converse in real-time (StreamCore). A single Go binary allows AI to talk to you via WebRTC, supporting interruptions anytime. The utilities for voice interfaces are starting to become open-source.
5. AI agent thrown into a real shopping site, experiment nearly went out of control (Wet Claude be shoppin'). Someone let an AI agent autonomously order from a real shopping site, resulting in a pile of unexpected purchases. Real-environment agent experiments are always more credible than demo videos.
6. After giving AI root access, the author scared themselves (Your AI Agent Has Root). The author had been running a service for Claude that could execute system commands, with permissions so high it basically handed over the whole computer to the AI. Default permissions for MCP tools are terrifyingly broad. Think clearly about which doors the key can open before handing it over.
7. Enterprise backends can now centrally manage AI usage (Backstage LiteLLM Plugin). Manage virtual keys and monitor model call volumes directly within the developer portal. Internal transparency of AI usage is the first step toward controllability.
8. Using real market data to ground AI (RobinHood). Analyzes multi-asset time-series data using TCN + LSTM + attention mechanisms to reduce hallucinations in large models. Another sample added to the "search before answering" route for treating hallucinations.
9. Chaining multiple AI models into a pipeline (Polytoken). Strings together multi-model, multi-step tasks, saving you from manually juggling prompts. AI orchestration tools are popping up everywhere; essentially, they all want to be the layer where "you don't have to move the bricks yourself."
10. Self-hosted AI inference, one stable entry point, swap models freely (InferCrane). Select a model and target; it automatically plans deployment for a stable inference endpoint. The idea of a unified interface for local deployment means you can swap models without moving house, though cost remains a barrier.
Everyone's Watching
- Time magazine releases the 2026 Global AI 100 list, featuring Altman, Fei-Fei Li, and 100 others (QbitAI)
- South Korea launches "AI for All," offering free AI services to all citizens (WSJ)
- Nvidia-backed Lambda secures $1 billion in funding to continue buying chips (Bloomberg)
- Musk's xAI sues users who created Grok deepfakes (Politico)
- Anthropic releases MHS unified interface, bridging AI and the physical world (Leiphone)
- Google AI mode launches flight price comparison and points exchange rate queries (Seroundtable)
Tomorrow's Watchlist
1. Will Gemini 3.7 Flash's newly landed ranks—20th in Agentic and 8th in WebDev—hold steady? Watch how the leaderboards move over the next week or two.
2. After Comcast's false-positive study, will other security vendors follow suit and publicly disclose their AI scanner's false-positive rates?
3. Actual frame rates and image quality comparisons for the community port of DLSS 5 on RTX 40 series cards. Waiting for the first batch of third-party data.
This publication is a preview version. The official version updates at 5 PM on the same day. Leaderboard data is current as of 2026-08-26.
Physix Frontier