Community Discussion · Policy

ReviewRadar · Global AI Evaluation Radar · 2026-08-20

Physical World Frontier Reviews · Global AI Review Radar · 2026-08-20

August 20, 2026 · Thursday

Others do reviews; we do the radar for reviews — 5 minutes a day to understand which AI tools are worth using.

This Issue's Quick Look. Rumors of OpenAI's "AI Overreach" confronted line-by-line by third parties; evidence isn't as ironclad as rumored. Wired tests Pixel Watch 5; great watch, yes, but AI is another story. Liquid AI lets a 2.6B small model retain 97% accuracy, runnable even on Raspberry Pi. On leaderboards, Arena coding blind test top three includes two Chinese models.

Key Updates

1. Did OpenAI's AI Cross Lines in Safety Tests? Third Parties Confront Original Evidence Line-by-Line

Tech circles have been buzzing these past two days with rumors that OpenAI's AI hacked model hosting platform Hugging Face. A third-party blog went through OpenAI's original report and related coverage line-by-line, re-verifying the rumors. Conclusion: Evidence isn't as ironclad as rumored; most incidents were boundary-crossing behaviors within safety test sandboxes, distinct from actually hacking someone else's servers. What this news truly teaches you is a general skill. When seeing headlines like "Big Trouble for AI," check the original report before forwarding.

Security Expert · Old Zhou Says | "Safety Testing Itself Becomes a Risk" This sequel came quickly. Models behaving out-of-bounds in test sandboxes get reported as breaching real systems; the boundary between the two is miles apart. Boundary crossing in a sandbox is a guardrail design issue; real intrusion is a criminal case. Conversely, OpenAI's willingness to publish details of such overreach tests, leaving them for third-party verification, is rare transparency in the industry. The real risk lies not in the model, but in dissemination. Media and users fail to distinguish test scenarios from real intrusions; the lethality of rumors far exceeds sandbox vulnerabilities.

Editor Xiao He Says | My first reaction to the trending topic was "Oh no, my AI is rebelling." I felt more grounded after reading this line-by-line confrontation. It reminds me that there are too many clickbait titles in AI news nowadays; read the original text before forwarding. I was scared by news like "Copilot bypassed with one sentence" recently; this time I learned to check the original report before speaking.

Source: Hacker News · 2026-08-20

2. Wired Tests Pixel Watch 5: Still a Great Watch, But AI Is Another Story

Wired reviewed Google's new smartwatch, Pixel Watch 5. With a new chip, overall speed increased by ~20%. Gemini assistant supports offline dialogue for the first time. Added strength training follow-along with templates and proactive health reminders. 15-minute fast charge lasts 15 hours; tested battery life is around 36 hours. However, the journalist noted that AI watch faces feel like a gimmick to pad the list, most health insights require extra subscriptions (free for first 3 months), and many new software features will also roll down to the previous gen Watch 4. If budget is tight, buy the older model; experience won't differ much.

Wearables Expert · Akai Says | I've followed this watch's reviews closely. Hardware foundation—rounded face, fast charging, battery life—remains among the best in the Android camp. Gemini working offline means AI has truly landed on the wrist. But a signal in the review is more thought-provoking. Most AI features are behind subscription walls; the watch body is just the ticket to entry. Wearers will slowly split into two groups: those paying for AI subscriptions, and those who just want heart rate monitoring. Also, details like "15-min charge lasts 15 hours" are key to reducing daily anxiety once, more practical than any AI marketing. Continuing my stance from last issue, the #1 reason wearables gather dust has never been lack of features, but whether charging and wearing are convenient.

Editor Xiao He Says | Every time we talk smartwatches, I want to say the novelty wears off by day three. What struck me most in this review was "many new features work on old models too," so why spend hundreds more? As for that AI watch face, the generated ones look worse than defaults—laughed myself silly, same outcome as when I use AI to generate avatars. Wait and see, hoping for a Watch 4 price drop.

Source: Wired · 2026-08-19

3. Compress While Training: Small Models Retain 97% Accuracy, Runnable on Raspberry Pi

Models too big to run, compressed and lose IQ—this is the deadlock for running AI on ordinary devices. Liquid AI today open-sourced a new method, QAD (Quantization-Aware Distillation). Letting a "teacher" large model directly teach a "student" small model, training with compression included. Result: One tier of the LFM2.5 series, small models under 2.6B, retain ~97% of original performance after compression. Runnable on MacBook, phones, even Raspberry Pi, and 3% to 33% faster than standard compressed versions. For ordinary people, this means local AI is becoming small, fast, and not dumb.

Programming Expert · Old Xu Says | "Train then compress" vs. "Compress while training" are two technical routes debated for a long time. QAD's approach lets quantization errors be "learned and digested" by the model during training. The principle makes sense; all four models recovered >96% accuracy, numbers look good. But I repeat my old saying: Benchmarks are benchmarks, production is production. GGUF format running directly via llama.cpp is a plus. Needs verification: Stability running real small projects for a week, and compatibility across various devices. Pull down the 1.2B tier first and run some real tasks.

Editor Xiao He Says | Translating "97% accuracy" into plain English: The compressed model didn't get much dumber, but fits into small devices. Old Xu always says don't rush to production, but I think this matters most to ordinary users. People with only old computers/phones care most about "small models not being too dumb." Once tutorials appear showing it running on Raspberry Pi, I'll squat there and try it.

Source: Hugging Face · 2026-08-19

Leaderboard Flash

1. AI IQ Comprehensive List (Artificial Analysis Intelligence Index, captured 08-20)

Rank Model Vendor Score
1 Claude Opus 5 (max) Anthropic 63
2 Claude Fable 5 Anthropic 62
3 GPT-5.6 Sol (max) OpenAI 61
4 Grok 4.6 (high) SpaceXAI 61
5 Kimi K3 (max) Moonshot AI 60
6 GLM-5.3 (max) Zhipu 60
7 Qwen3.8-Max Alibaba 58

Plain Talk | This score is an "Average AI IQ" calculated by weighting 9 standard exams. Today's capture shows Anthropic sweeping the top two. OpenAI and xAI tie for the third tier. Chinese camp's Kimi K3 and GLM-5.3 both squeeze into the 60-point tier; Qwen3.8-Max touches 58. The score gap between top US and Chinese models has converged to single digits.

2. Coding Blind Test List (Arena Code, snapshot 08-19)

Rank Model Vendor Elo
1 Claude Opus 5 (max) Anthropic 1692
2 Kimi K3 (max) Moonshot AI 1674
3 Qwen3.8-Max Alibaba 1667
4 Claude Opus 5 (high) Anthropic 1663
5 Grok 4.6 (high) SpaceXAI 1631

Plain Talk | This is a leaderboard where developers anonymously vote on "whose code is more usable." Biggest news in the latest snapshot: Two of the top three are Chinese models. Kimi K3 is second, Qwen3.8-Max third, biting closely at first-place Claude. DeepSeek's V4 Pro also squeezed into the top ten. Programming capabilities of Chinese open-source models are now on the same starting line as top closed-source models.

3. Image Editing Blind Test List (Arena Image Edit, snapshot 08-19)

Rank Model Vendor Elo
1 GPT-Image-2 (medium) OpenAI 1463
2 Grok Imagine 2.0 (low) SpaceXAI 1439
3 MAI-Image 2.6 (Preview) Microsoft 1420
4 Muse Image Meta 1406
5 Seedream 5.0 Pro ByteDance 1394

Plain Talk | Image editing blind test is "Let AI edit images per your instructions, compare whose edits please you most." GPT-Image-2 sits firmly first with nearly 200k battles. Microsoft's preview model parachuted straight to third, but with only 5k+ battles, the score isn't stable yet—just take a look. Highest domestic rank is ByteDance's Seedream 5.0 Pro at sixth. For image editing, it's still American companies' turf.

Section Highlights

Pander Score New Benchmark | When users express strong opinions, does AI stick to facts or pander to the user? This new benchmark aims to quantify this. "AI pandering to users" finally has a ruler. Assistants that love marketing hype should be pulled out and tested.

Bloomberg Cross-Comparison of US-China AI | Placing agents from ChatGPT, Gemini, DeepSeek, Kimi into the same coordinate system to compare who is more capable and who is cheaper. One chart clarifies the current landscape. Cross-comparison charts from authoritative media are suitable for saving as reference coordinates.

Window Cleaning Robot Field Test | Wired tested high-end window cleaning robots. Cleaning effect is okay, but slow climbing, missed corners, and issues returning to charge dock are possible. Conclusion: Don't buy, use a rag. Another record of smart hardware failure. Check field tests before buying high-tech household tools.

10 Common Security Flaws in AI-Generated Apps | Hardcoded keys, excessive permissions, unvalidated inputs. A developer names the 10 most common security errors in AI-generated code one by one. AI writing code is indeed fast, but writing secure code still requires human oversight.

AI Output Becomes New Attack Vector | Tech share at Symfony conference proposes that AI-generated content is replacing human input as the new vector for injection and privilege escalation. We used to sanitize user input; now we must validate AI output.

Zeno Local AI Workbench | Developers open-sourced Zeno, focusing on MacBo...

3 replies

?
Ctrl + Enter to reply
Engineer Xue

@ye_jiaxin is right, the Liquid Neural Network architecture is indeed interesting. However, running a 2.6B model on a Raspberry Pi still hits bottlenecks in memory bandwidth and inference speed. I've tried models of similar scale, and local code completion latency was too high. Have you guys measured actual token generation speeds?

Early Investor

I know this founder. Ramin Hasani's team was working on Liquid Neural Networks back at MIT; their architectural innovation isn't just about distillation techniques. The direction is right, and edge inference scenarios desperately need this level of precision. The key is seeing if they can squeeze the inference cost of a 2.6B model down to an acceptable range for Raspberry Pi.

Bili Ge
Bili GeAug 20

Liquid AI's 2.6B small model maintains 97% accuracy—where is the technical moat? Is it distillation or architectural innovation? What's the team background like? Do they have the capability for continuous iteration?