Community Discussion · Policy

ReviewRadar · 2026-08-27

Physix Frontier Review · 2026-08-27

This issue's theme: AI writing code—speed is fine, but stability must keep up.

Key Updates

1. Official Release of "Handbook for Software Development in the AI Era"

Anthropic officially released a handbook on AI-native software development workflows. The core idea: break large tasks into small steps; for each step, let AI propose a plan first, get human approval, and follow with testing. The official claim is that some teams complete a week's worth of work in a day, but if the process doesn't change, rework and incidents will also double.

Programming expert Old Xu says small-step commits, test-first, and humans reviewing code are consistent with his 15-year-old rules. However, the handbook was written by Anthropic itself, revolving around its own models. Whether it can be implemented depends on waiting for other teams to try it out.

Editor Xiao He says verify AI-written code before trusting it. Like checking a package upon arrival for damage, the handbook teaches how to unpack and inspect.

2. Installing a Security Checkpoint for AI Coding Agents

The open-source tool Grith monitors every action of AI coding agents at the system level—reading files, deleting things, connecting to networks, executing commands. Anything not approved is blocked, and past actions can be replayed.

Security expert Old Zhou says permissions should be set to the minimum. This turns "environment controls everything" into an off-the-shelf tool, hitting the nail on the head. The cost is speed loss as every action passes through the gate. Whether there are loopholes requires real-world project logs; don't just look at the intro docs.

Editor Xiao He says when AI works in my house, I can't give it all the keys; every door opened needs my permission. Don't rush to authorize important accounts.

3. AI Agent Works for 14 Months, Humans Become Supervisors

An engineer wrote about their 14-month experience directing AI agents to write code. The prerequisite for stable AI delivery is breaking tasks down small enough and having humans verify each step. Humans haven't become idle; they've just changed jobs.

Programming expert Old Xu says this aligns with what he always preaches: "As AI output increases, verification must keep up." Be most wary of "fully automated ATM" marketing pitches; generating code doesn't mean it can go live.

Editor Xiao He says AI helps people work, but humans still check quality. Like driving with navigation, AI finds the route, but you still hold the steering wheel.

Leaderboard News

Alibaba Open-Sources Qwen3.8-Flash (Released Aug 26)

Total 125 billion parameters (brain cells), waking up only 6% during work. It beats Claude Opus 4.6 in several exams: leading by 9.1 points in coding, 22.5 points in phone operation, and 25.1 points in solving math problems from images. Price is only 3% of Opus 4.6, and training costs dropped 90% compared to the previous generation. These are vendor-published numbers; per usual practice, wait for third-party re-testing.

Human Blind Test Leaderboard (Data as of 2026-08-26, Not Updated Yet)

Real humans vote and score AI responses to rank them. Four of the top five spots are held by Claude. Meta's Muse Spark 1.2, ranked 4th, has only had over 3000 matches; sample size is too small, so don't take the score seriously yet.

Agent Leaderboard (Data as of 2026-08-26, Not Updated Yet)

Tests whether AI working for you is reliable, if it can recover after mistakes, and if it follows instructions. Among the top ten, domestic Kimi K3 ranks 6th, the highest among Chinese models.

Section Highlights

  • A programmer with 25 years of experience reflects that AI makes him code faster but makes him miss the joy of "playing"
  • Google uses AI assistants to rewrite old C/C++ code into Rust, which is harder to have vulnerabilities, with human verification
  • Devx goes open source; autonomous AI coding agents can run on Android phones with Termux
  • GitHub official tutorial: Let Copilot automatically handle code reviews submitted by bots
  • Wired covers the humanoid robot games live: running faster than Bolt, tweezers skills even more impressive
  • Physlint goes open source, performing health checks on robot training data to detect dirty data and prevent fraud
  • Waymo has driven 200 million miles, publishing 10 lessons learned on autonomous driving
  • thevibes.dev launches, scoring and reviewing AI models—a Dianping for models
  • Running AI models at home requires patience first; installing environments and weights is physical labor
  • Neoswarm goes open source, commanding a swarm of AI agents to work within the editor

Everyone Is Watching

Nvidia's quarterly revenue approaches $100 billion; Zhipu's "Niulai" is actually GLM's first native multimodal model; OpenAI publishes full report on AI agent intrusion into Hugging Face; Tiangong Ultra breaks 100m record again in 8.64 seconds; World Humanoid Robot Games conclude, Unitree disciple defeats master to win first gold in free combat.

Tomorrow's Focus

① Qwen3.8-Flash weights are open-sourced and online; when will third-party re-tests come out?

② Arena snapshot hasn't updated for two days; will a new leaderboard appear tomorrow?

③ Multimodal capability tests after Zhipu "Niulai" weights release

Physix Frontier Review Editorial Team · Issue No. 029 · Preview Version

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts