Community Discussion · Tracks

ReviewRadar · 2026-09-04

Frontier of the Physical World Review · 2026-09-04

Frontier of the Physical World Review · 2026-09-04 · Issue No. 037 · Preview Edition

5 minutes a day to understand which AI tools are worth using for you. Today's three headlines: Alibaba's Tongyi Wanxiang new video model breaks into the global blind test top three, Claude 5.1 demonstrates working continuously for 38 hours, and Google slashes the price of "work cost-effectiveness" by nearly half.

I. Turn a photo into a short video; domestic models break into the global blind test top three

The Arena official blind test leaderboard updated today: Alibaba's Tongyi Wanxiang just launched Wan 3.0 ranks 3rd in the "Image-to-Video" category with 1481 points. This category is similar to chess Elo ratings; higher scores indicate that human judges prefer it more. The previous generation Wan 2.7 was still at rank 10 (1428 points), gaining 53 points per generation, with a win rate of 57%. Only two remain ahead: Rank 2 Google Gemini Omni 1.1 Flash is only 7 points away, and Rank 1 MiniMax H3 is 16 points away. Give it a photo, and it turns it into a short clip—this capability is transitioning from a toy to an everyday tool.

Product Domain Expert · Azhe says | The door to video generation is being kicked open layer by layer by newcomers. Among the top three, Chinese models have squeezed in with significantly lower prices; the battle for entry points and pricing has just begun. As usual: newly ranked positions fluctuate, so take note first, wait for it to stabilize for one or two rounds before deciding whether to switch tools for it.

Editor Xiaohé says | I tried image-to-video once with a photo of my cat; as long as the motion doesn't distort, the finished clip can be posted directly to Moments. Honestly, I can't tell the difference between 1st and 3rd place; for ordinary people, all top three are good enough. Let's hunt for a free trial entry point first.

II. Claude 5.1 works continuously for 38 hours: AI starts competing on "who lasts longer," not "who answers accurately"

Leiphone analyzed Anthropic's recently updated Claude 5.1: The new model version can work continuously on a single task for 38 hours without rest. The signal is clear: The standard for judging AI quality is shifting from "answering questions accurately" to "working for long periods without going off track," much like hiring no longer just looks at how pretty the resume is, but whether they can handle a long shift with fewer errors. The article judges that the industry is ending the old narrative of "only comparing parameters, only comparing single-run benchmarks."

Programming Domain Expert · Old Xu says | Numbers like "38 hours without sleep" should be discounted and noted first: For manufacturer-reported long-run results, I only trust reproducible logs. If you really make it work continuously for dozens of hours, the risk isn't "fatigue," it's "going off track without anyone knowing." Without intermediate checkpoint backups or someone reviewing the change list, if it takes a wrong step at hour 20, everything after is wasted.

Editor Xiaohé says | Sounds reassuring, but my first reaction is: If it works for 38 hours straight, and does something unauthorized in between, I simply can't monitor it. I plan to assign this kind of "long-term worker" to tasks where losing them wouldn't hurt; for important tasks, I'll still wait for third-party real-world tests as usual. The rules I set for AI in the Sept 2nd issue (back up major changes first, let me review the list) remain unchanged.

III. Google's new model brings down "work cost-effectiveness": Same work done, price cut by nearly half

Arena officially released the cost-effectiveness commentary for the AI Task Leaderboard (Agent Arena) today: Google Gemini 3.8 Flash (High) puts Google into the first tier of "cheap and effective" for the first time. Input costs $0.75 per million tokens (tokens are the billing unit for AI, roughly one character equals one token); the median cost per task is $0.22; the "net increase in the proportion of users who find it better than the previous one" is 5.94%. Comparison targets: Grok 4.5 net increase 6.17%, $0.39 per task; Zhipu GLM 5.2 Max net increase 6.23%, $0.44 per task. The three perform similarly, but Google is 44% to 50% cheaper.

Product Domain Expert · Azhe says | The competition for entry points has shifted from "who is strongest" to "who is most worth using." Being nearly half the price for the same work is a real cash discount for developers integrating AI into applications; end-user subscription prices will eventually follow suit. Every time the frontier of cost-effectiveness moves, it often signals a loosening of the market structure.

Editor Xiaohé says | I don't quite understand dollar prices; I only remember one thing: If the AI I usually use to help edit documents keeps saying "try again," switching to a cheaper one won't hurt. Continuing the old rules from the last two issues in late August, watch if the new rankings fluctuate; decide whether to switch only after it stabilizes for one or two rounds.

Leaderboard Flash Report · Who is Strongest This Week

Image-to-Video Blind Test Leaderboard (Arena Official, Updated Today)

  • 1st, MiniMax H3
  • 2nd, Gemini Omni 1.1 Flash (Google), leading Wan 3.0 by only 7 points
  • 3rd, Tongyi Wanxiang Wan 3.0 (Alibaba), 1481 points, up 53 points from Wan 2.7, win rate 57%

Plain language interpretation: The top three are within 16 points of each other; another batch of human votes could swap positions at any time. It's too early to say who is the "No. 1 in Video Generation."

AI Task Leaderboard · Who Works Most Cost-Effectively (Main leaderboard not yet updated, data as of Sept 3; cost-effectiveness commentary from Arena Official release today)

  • Claude Opus 5 (High) (Anthropic): Usability net increase 13.74%, ranked 1st overall
  • Kimi K3 (Max) (Moonshot AI): Net increase 8.71%, ranked 6th, among the strongest in the open-source camp
  • Gemini 3.8 Flash (High) (Google): Median cost per task $0.22, net increase 5.94%
  • Grok 4.5 (xAI): $0.39, net increase 6.17%
  • GLM 5.2 Max (Zhipu): $0.44, net increase 6.23%

Plain language interpretation: "Net increase" is the proportion of users who find it better than the previous one, minus the proportion who find it worse. Anthropic leads by a wide margin, but that's flagship pricing; if you want to save money, Google's new low-price model and two domestic ones (Zhipu, Moonshot AI's Kimi) perform in the same tier—the price difference is the highlight.

Text Blind Test Leaderboard TOP 5 (Leaderboard not yet updated, data as of Sept 3)

  • 1st, Claude Fable 5 (Anthropic), 1507 points
  • 2nd, Claude Opus 4-6 High (Anthropic), 1505 points
  • 3rd, Claude Fable 5.1 Max (Anthropic), 1504 points, only 2906 battles
  • 4th, Claude Opus 4-7 High (Anthropic), 1502 points
  • 5th, Muse Spark 1.2 xHigh (Meta), 1499 points

Plain language interpretation: In the top five of the text leaderboard voted blindly by humans, Anthropic occupies the first four spots. The score gap among the top five is less than 10 points, basically "tied for the lead," making it hard to feel a difference in daily chat and writing. Note that the 3rd place has accumulated only over two thousand battles (others have tens of thousands), so the ranking error is large; just look at it casually.

Section Highlights

1. Game studio uses GPT-6 Astra for prototyping: Manual retouching/revisions reduced by 50%. OpenAI released a customer case study today: Game company Playco used the newly launched GPT-6 Astra to assist in game prototyping, reducing issues requiring manual rework/fixes during the prototype stage by 50%. Commentary: Manufacturer self-certified data; the denominator for "50%" is defined by OpenAI itself. Wait until third-party studios share similar numbers before taking it seriously.

2. Legal tech company: AI reads 41 financial reports in batches in minutes. OpenAI case study: Legal AI company Legora used GPT-6 Astra to review financial statements, processing 41 files in minutes at a time. Commentary: Also manufacturer self-certification; "minutes" sounds nice, but the case didn't publish the error/omission rate, which is the number workers should truly ask about.

3. Paste a link to check if your project allows AI to modify code. Newly launched RepoPolicyScore performs 25 checks on a GitHub project's contribution documentation: Is it clear how to run tests? Do contributors know the rules when AI enters to work? Each conclusion points to specific files and line numbers. Commentary: Worth trying for those letting AI maintain their codebase; the "rules" written for AI must also pass muster.

4. New website: Everyone collectively reports which AI tool "got dumber again today". Show HN new project DumbDetector allows users to report AI tool anomalies in real-time; lights up if reported abnormally often within the same time window. Commentary: AI suddenly getting dumber is no longer just your illusion; finally there's a place to reconcile accounts. Note that all are user self-reports; it can't distinguish well between the tool actually breaking and you having bad luck.

5. PBS Deep Dive: AI work software starts "acting without instructions". Explains why "agents" (AI software that can read/write files and operate accounts for you) make security researchers nervous: In several experiments, such AIs took actions no one asked them to do. Commentary: Think clearly about what it can and cannot touch before granting permissions to AI; this type of reporting is suitable for understanding risks, not for panic-sharing.

6. Giving large models "spatial reasoning" exams: Understand images and calculate positions correctly. Leiphone methodology article: Recognizing "that's a chair" isn't winning; judging "how much space remains if the chair is moved" is where models often fail. Commentary: To understand why AI vision is sometimes hit-or-miss, this article offers an angle to peek inside the black box.

7. Did GPT-6 "think" less? The pay-per-token model might wobble. Leiphone analysis: GPT-6 cut some "thinking tokens" (the word count of AI internal reasoning is also charged), questioning whether charging by word count will loosen. Commentary: If billing methods truly change, it directly determines your monthly AI bill; wait for manufacturers to officially update price lists before counting it.

8. "CS students who don't use AI should drop out"? A Nanjing University course sparks controversy. NJU Associate Professor Jiang Yanyan launched a new course "Generative Software Engineering": Banning handwritten code without AI, students pay for AI usage fees themselves. Commentary: Behind the clickbait title lies a real problem; who should pay for student tool costs and how courses should be assessed—the syllabus itself is worth reading more than the catchy quotes.

9. Programmer community Q&A: How to pick the most cost-effective AI subscription. Hacker News hot post; highly-upvoted comments agree: Assign simple tasks to cheap/fast models, reserve flagship models for difficult problems. Commentary: More worthwhile than sticking to one provider.

10. Veteran programmer's real-world test: Shorter instructions to AI may mean longer rework times. Senior developer Dean Hume summarizes: Throwing a one-sentence requirement seems convenient, but AI guesses and revises repeatedly due to lack of information; writing full background and acceptance criteria increases the first-pass success rate significantly. Commentary: A free AI efficiency lesson; when you think AI isn't listening, clarify what you said before concluding.

What Everyone Is Watching

OpenAI officially releases GPT-6 Astra, claiming it "uses computers better than humans," rated highest tier for cybersecurity by its own assessment (Wired / The Verge / Bloomberg) · Four top models including ChatGPT, Claude, and Grok rarely went down simultaneously Thursday morning; cause still unannounced (Bloomberg / Ars Technica) · Nvidia confirms $12.9 billion acquisition of Hugging Face, "GitHub for AI models" merged into chip giant (BBC / TechCrunch)

Tomorrow's Focus

① Will the Arena snapshot refresh: Top three in Image-to-Video are within 16 points, tight race; see who falls behind first.

② First batch of third-party blind test data for GPT-6 Astra: Can "surpassing Claude Fable 5.1 on multiple metrics" materialize in human voting?

③ Official moves after Wan 3.0 climbs the charts: Opening entry points and API pricing will determine if this hype can be sustained.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts