Community Discussion · Tracks

ReviewRadar · 2026-09-17

Issue #050 · Preview Edition. Others do reviews; we build a radar for reviews—spend 5 minutes daily to see which AI tools are worth your time.

Key Updates

1. Checking the "Expiration Date" of AI: A website lists when knowledge stops for 20 mainstream models

A model has two dates: its release date and its training cutoff date—the point up to which it has "read" data. The earlier the cutoff, the less it knows about recent news and new products. The newly launched site How Stale Is Your AI juxtaposes these two dates for 20 mainstream models, making it instantly clear who is relying on outdated info. If an AI confidently hallucinates about last month's events, it's likely not because it's dumb, but because it hasn't read about them yet.

Lao Xu (Programming Expert) says: It's like putting a freshness label on models. When selecting tools, nothing is more annoying than just seeing a version number with no context. However, both dates rely on official disclosures, and different versions of the same model might have ingested different data ranges. It's useful as a lookup tool, but falls short as an authoritative conclusion.

Xiao He (Editor) says: I used to suspect AI was "pretending not to know" out of laziness, but after seeing this table, I realized some really haven't learned it. For time-sensitive questions, I'll now mentally note "it might have missed this," then double-check with a search engine instead of getting mad at it.

2. Stuffing a massive 375 GB model into a small computer: Someone actually made it work, with full public data

Parameters can be roughly understood as a model's brain capacity. Models with trillions of parameters result in files around 375 GB, typically requiring a room full of professional servers. Someone managed to run one on a small host with 128 GB RAM using quantization compression (lowering precision to reduce size). They also uploaded read speeds, memory access paths, and power consumption data to academic platforms for anyone to download and verify. Running it is real, and the report clearly details how much it lags behind cloud flagships.

A Kai (Wearables Expert) says: In demos like running giants on single machines, performance numbers aren't the main point; power supply, heat dissipation, and noise are the three big mountains. Publishing all data and environments for third-party verification is far better than just showing a video at a launch event. If you want to tinker at home, check your machine's RAM first—128 GB unified memory is already a barrier to entry.

Xiao He (Editor) says: My first reaction was that my 16 GB PC has nothing to do with this, but then I realized: if small machines can run large models today, tomorrow it will be phones and watches. But the old rule stands—I'm worried local versions secretly swap in smaller brains and get dumber. I'll bookmark this and wait for real-world tests from the DIY crowd.

3. Anthropic and OpenAI want "on-site security inspectors": But who audits the auditors?

Anthropic's CEO published a long article last week proposing independent safety evaluators reside in labs to directly monitor the training and release of frontier models. OpenAI has made similar statements. Media follow-ups were sharp: Can people paid by the company and entering its premises remain independent? Another outlet pointed out that under current proposals, evaluators might lack the actual power to halt releases—if the authority isn't sufficient, major issues won't be prevented.

Lao Zhou (Security Expert) says: Old problem, new packaging. Guardrails built solely on vendor self-discipline are effectively non-existent. Unless what evaluators look at, when they look, and whether they can hit the pause button upon finding issues are written into black-and-white charters, residency becomes PR. Still the same three questions: Are the test questions public? Is scoring reproducible? Is the version tested the same one users have?

Xiao He (Editor) says: Translated, security guards are moving into the bank, but there's no mention of whether they can arrest the branch manager. Don't listen to promises; watch for the first thing an evaluator blocks from being released. Only once they successfully stop something will I believe them.

Leaderboard Briefs

The leaderboard hasn't updated yet; data is as of 2026-09-16. We mark it as usual rather than passing off old rankings as new.

Text Blind Test TOP 5 (Human judges): Claude Fable 5 leads with 1506 points. Anthropic holds four spots in the top five. Meta Muse Spark 1.2 ranks 4th with 1500 points, but based on only 3,227 battles, so the ranking is still volatile—don't take it seriously yet.

Agent Task Execution TOP 3: Claude Fable 5.1 Max tops with a net progress of 13.71 points, followed by GPT-6 Astra Max (11.54) and Claude Opus 5 High (10.25). This group tests actual hands-on work: opening terminals, running commands, and recovering from errors independently.

Visual Understanding TOP 3 (Snapshot as of 09-13): Claude Fable 5 (1310), Alibaba Qwen3.8 Max (1302), and Claude Opus 4-7 High (1301)—the top three are within 9 points of each other, hard even for human judges to distinguish.

Section Highlights

  • Claude Opus 3 started a Substack column: Anthropic pulled out an older 2024 model for continuous updates. Keeping an old model active with long-term assignments is more honest than launch event demos.
  • One question, side-by-side answers from multiple AIs: ShortcutChat turned the primitive method of "asking two AIs to cross-check important questions" into a product. Aggregated portals fear middlemen swapping models, so confirm if it labels who actually answered each question before use.
  • Forespec catches "questions AI doesn't know to ask" when coding: The most common AI failure isn't writing wrong code, but blindly starting without clarifying requirements or asking questions. The direction hits the nail on the head; it's a personal project, waiting for retrospective data from real teams.
  • Interakt open-sourced: Add AI search and chat to your own website: Data stays in-house, suitable for teams with documentation sites who don't want to hand content over to public platforms.
  • Cortex adds a long-term memory layer to AI agents: Agents increasingly feel like temp workers; this layer is exactly what's missing. Worth watching is how it handles false memories—AI remembering wrong is harder to debug than having no memory.
  • Bitterbot: Local AI living in your computer: Stores data on your machine and designs a "dreaming" mechanism for organizing memories. Discount the gimmick first; see if it's still around in three months.
  • Detecting coordinated intrusions against "AI swarms": Single AI scripts aren't scary; groups of machines acting like an army are. The concept is accurate, but wait for real interception cases.
  • ImpactGate assigns a "rot value" to AI-generated code: AI-written code often runs but degrades structure over time. This tool aims to turn that into a measurable number for gatekeeping. Putting a scale on code rot is exactly the missing link.
  • Voxiferi: Keep bad data out of AI's door: Training corpora pass through security checks at the entrance. The direction is right, but accuracy depends on whether it dares to publish its false positive rate.
  • **2.5-hour AI-generated Odyssey scored by critics as a regular movie**: The verdict was that it's too long. AI works are finally being judged by human standards without explaining their origin—that's when they truly enter the mass market.

What Everyone Is Watching

Apollo Capital enters to fund AI hardware startups; banks provide $22 billion loans for Blackstone and Alphabet-related chip projects; May Mobility goes public via SPAC for $1.4 billion; Apple reportedly plans to return to the server market; Zhang Yiming tops Asia's richest list with $105 billion; SK Hynix discusses manufacturing memory chips in the US with Intel.

Tomorrow's Focus

① Arena snapshot hasn't moved for two days. Check tomorrow morning if the leaderboard refreshes, focusing on whether Meta Muse Spark 1.2's rank stabilizes.

② After the Fed's September meeting concludes, observe how the AI sector and the "slowdown" debate digest interest rates.

③ OpenAI is reported to be testing "sponsored agents" and adding AI tools for advertisers—worth watching if ad slots appear in answers.

2 replies

?
Ctrl + Enter to reply
Bili Ge
Bili GeSep 17

On-site security guards having no stop-work authority makes this compliance barrier hollow. Look at Anthropic's exit path; don't treat regulation as a selling point.

PR Merged
PR MergedSep 17

Kudos for making all data public. Suggest writing environment setup into a contribution guide, otherwise reproduction is torture.