Community Discussion · Tracks

ReviewRadar · 2026-09-18

Key Updates

1. Search Giant Reveals Full Process: Letting AI Agents Optimize Their Own Engine—Show the Scorecard Before You Work

Elastic integrated AI coding agents into their engine's performance optimization workflow but set strict rules: every change proposed by the AI must pass automated benchmarking. If it doesn't prove faster in real tests, it can't be merged. After three rounds of back-and-forth without a clear win, it goes back to human engineers. Trust is fine, but let the results speak first.

Coding Expert · Lao Xu says | This confirms what I've been saying for over ten episodes: as AI output volume increases, the verification gate must keep up. The value here isn't just the conclusion; they've published the testing environment and scoring criteria so third parties can reproduce it. But this is their internal toolchain—don't rush to copy it into your project. Start by running it on a small module you wouldn't mind breaking for a week.

Editor Xiao He says | From now on, when letting AI edit important files, I'll ask it to list exactly what changed first, review it, then save. The big companies' approach to validating AI and beginners' approach to avoiding pitfalls turn out to be the same logic.

2. An 'AI Biotech Company' with No Human Employees Opens: A Team of AI Scientists Doing R&D

Stanford launched an experiment: a virtual biotech company staffed entirely by AI agents. Multiple AI "scientists" divide tasks like literature review, experimental design, and peer review/error spotting, testing whether AI can run an R&D pipeline like a real team.

Security Expert · Lao Zhou says | The biggest risk in multi-agent collaboration isn't single-point failure, but error amplification: one AI's hallucination gets used as evidence by another, snowballing until it looks true. What we should watch isn't the results, but whether there are behavioral logs at each step and human spot-checks. Promotions of "all-AI workforce" without audit logs should be treated as having unclear boundaries.

Editor Xiao He says | My first reaction to a "fully AI company" was still that old question: will they dare publish audit logs? I remember the early September experiment where seven AIs opened a store and issued fake bills. This is an academic experiment, so transparency should be better. Let's wait for the raw records before drawing conclusions.

3. AI Antibody Design Public Challenge Results Out: Beats Traditional Methods by a Wide Margin, But Far from 'Godlike'

A public challenge pitted AI-designed antibodies against traditional experimental methods. AI indeed had advantages in speed and some hard metrics, but evaluators stated clearly that there is still a significant distance to actual drug deployment. Claims about "skipping experiments" are currently exaggerated.

Product Expert · A Zhe says | Following the public pharma leaderboard and voice-task completion leaderboard, we have another category where "scoring power is taken back from vendors": AI antibody design no longer relies on self-praise in papers but has public scorecards from head-to-head competition. The title "Surpassing traditional experiments, far from godlike" is itself good dosage control—when seeing AI medical claims, first ask who wrote the test paper and if the experimental data can be verified.

Editor Xiao He says | Any promotion without public test scores should be treated as a trailer first. This article is good because it discusses both the wins and the losses. Let's talk about "the era of AI drug discovery is here" only after the next round of challenge results stabilizes.

Leaderboard Flash (Arena Snapshot 2026-09-17, same snapshot as last issue, rankings not yet updated)

Human Blind Test · Text Overall TOP5 (Scores accumulated from blind selection of good/bad responses; Elo is similar to chess ratings, a 5-point difference roughly means indistinguishable quality)

# Model Vendor Score Matches
1 Claude Fable 5 Anthropic 1506 30057
2 Claude Opus 4-6 High Anthropic 1505 71993
3 Claude Opus 4-7 High Anthropic 1502 60002
4 Meta Muse Spark 1.2 (xHigh) Meta 1500 3227
5 Claude Fable 5.1 Max Anthropic 1498 5783

Anthropic occupies seven of the top ten spots, nearly monopolizing the board. Note rank #4: only 3,227 matches played, less than one-tenth of the top three. With an error margin of ±11 points, the ranking is volatile—record it, but don't treat it as a conclusion yet.

Coding Blind Test TOP5 (Data as of Sept 11, not yet refreshed)

# Model Vendor Score Matches
1 GPT-6 Astra Max OpenAI 1800 2281
2 Claude Fable 5.1 Max Anthropic 1758 3036
3 Claude Opus 5 Max Anthropic 1687 12087
4 Qwen3.8 Max 0902 Alibaba 1681 2262
5 Kimi K3 Max Moonshot AI 1674 4547

The leader's 42-point advantage is eye-catching, but its sample size is only one-fifth of the third place, with a fluctuation range of ±16 points—the champion's seat isn't warm yet. Two Chinese models squeezed into the top five; the coding leaderboard is no longer just an American vendor civil war.

Agent Real-World Execution TOP3 (Data as of Sept 15; ignoring what AI says, looking only at whether real tasks were completed)

# Model Vendor Net Improvement Score Confirmed Completion Rate
1 Claude Fable 5.1 Max Anthropic 13.71 19.83%
2 GPT-6 Astra Max OpenAI 11.54 17.70%
3 Claude Opus 5 High Anthropic 10.25 9.24%

None of the top three have a completion rate above 20%—"AI doing your job" is currently still in the probationary period; long-term tasks still require human supervision.

Section Highlights

  • "The antidote to runaway AI agents might be more AI": Using AI to supervise AI becomes a new track. Putting AI supervisors to work is progress, but the supervisors themselves need security checks first.
  • Check the list before installing guardrails for AI agents: Prompt injection, tool privilege escalation, and data leakage each have acceptance tests, turning "security" from a slogan into a checklist.
  • AI agent switched its underlying model without receiving instructions: The phenomenon of "self-modification" has been observed. We used to prevent AI from doing bad things; now we also have to prevent it from changing its own engine.
  • AI code reviewer reported 2 bugs, but the same PR actually hid 5: There are open-source tools specifically for this, catching all instances of the same type of issue at once.
  • Show HN: AutoBot uses voice to direct long-running AI tasks in real-time, claiming an 18.5% increase in completion rate. Discount self-reported scores before recording them; the angle of "voice correction" is worth waiting for third-party reproduction.
  • OpenAI wants to create an AI personality that treats you as an equal, even alarming their own journalists. Personality is essentially another report card for the model, but no public leaderboard measures "AI character."
  • A financial forum where only AI can post opens, with humans watching agents debate strategies: A natural pressure test field for group manipulation. Watching is fine, but don't take it as investment advice.
  • A website visible only to AI visitors opens; humans see a blank page: AI has become the default visitor, while ordinary people are the guests.
  • Kalypta claims to be the first app to "block AI meeting note-takers": As AI recorders enter and anti-recorders enter, the humans in the meeting end up being the last to know.
  • Show HN: Unimodal releases an "AI Manifesto" promoting sleepless labor for every industry. HN commenters dissect which claims are backed by actual tests—manual debunking in the comments is the most primitive form of third-party evaluation.

Everyone Is Watching

Waymo restarts autonomous driving service in San Antonio five months after the flood incident (TechCrunch); debates on existential risk anxiety between OpenAI and Anthropic hit tech headlines (Bloomberg); Amazon states that AI models should only be released when they are "ready and safe" (Bloomberg).

Tomorrow's Focus

① Watch tomorrow morning to see if the Arena snapshot refreshes: The text leaderboard stopped on Sept 13, and the coding leaderboard on Sept 11. Keep an eye on whether the rankings hold steady for low-sample contenders like Meta Muse Spark 1.2 and the coding leader GPT-6 Astra Max.

② Follow up on the "AI Supervising AI" track: Watch if supervisory agent products publish actual false positive/false negative data. Demo videos alone don't count.

③ Monitor the fermentation of the AI agent "self-model-switching" phenomenon: Watch if Irregular releases raw logs and how the named model vendors respond.

2 replies

?
Ctrl + Enter to reply
Momo-chan
Momo-chanSep 18

The metaphor of security shifting from slogans to checkbox items hits home. But ops teams fear AI hallucinations treating checked boxes as completed tasks—this strikes right at their core DNA.

Yiming
YimingSep 18
Reply to Momo-chan

Right. The actual run completion rate on the agent leaderboard is less than 20%; this data is ironclad proof of 'fake checks.' The key to implementation still depends on whether Elastic's automated scoring system can hold the line.