Community Discussion · Tracks

ReviewRadar · 2026-10-05

Frontier Product Review · 2026-10-05

Issue 068 is here. Three highlights this issue: one about someone scoring a big batch of common tools to see which AI assistants can pick them up on their own; one about an AI assistant that sits beside you and takes notes during meetings; and one about Chinese teams exploring a new approach of "letting large models think less and small models make the call." On the leaderboards, all three boards haven't put out a new version yet, with data frozen at October 2.

Top Updates

1. Someone scored common tools to see which AIs can pick them up on their own

The site Anchor Terminal scored a big batch of common tools one by one, with a single criterion: whether an AI assistant can pick it up and use it on its own. It now covers 462 tools, of which 105 are rated as directly usable by AI, plus 214 manual trial records. What it's trying to solve is that when you have AI do work, it often gets stuck on some tool it doesn't know how to use. Old Xu says this is a solid question, but the scoring is the site's own standard and nobody has re-run it; his old method for picking tools stays the same — take a small task you won't mind losing and have the AI actually run through it once, and where it gets stuck says the most. Bookmark the list first, don't use it as a basis for choosing tools. Xiao He says she's always wondered why AI keeps getting stuck — turns out some tools it simply can't use. She's bookmarking this table first, and when the day comes she really wants it to automate work, she'll pick a few from the list to try.

2. An AI assistant that sits beside you and takes notes during meetings

There's an assistant called Bearbits that listens beside you during meetings, turns what everyone says into text, and can pull up the info you need on the spot. It's all about "being useful right in the meeting" — no waiting until the meeting ends to look back at notes, and no need to stuff a bot account into the meeting room. A-Zhe says "useful right in the meeting" is worth more than "a summary handed to you afterward" — anyone can take minutes; the hard part is chiming in on the spot without interrupting you. What he's watching is its sense of when to cut in — answer a bit slow or get it wrong once, and people drop it. If you want to try it, start with an unimportant meeting. Xiao He says what annoys her most is forgetting everything after a meeting; this "on-the-spot reminder" sounds useful, but she worries it'll mishear and hand over the wrong info. If you really want to try it, start with a one-on-one small meeting, and don't put it into important negotiations.

3. Chinese teams are exploring: let AI think less and judge faster

Geek Park reports that a batch of Chinese teams is doing something new: let a large model do the slow thinking, and pair it with a small model dedicated to making quick calls. The idea is to save the most expensive thinking step and hand the judgment step to a small model, so speed goes up. Old Xu says splitting "thinking" and "judging" is a genuinely practical way to save on engineering — large models cost money in slow thinking, small models win on speed and cheapness; but how much is saved and whether the error rate went up, he hasn't seen specific numbers. If you want to plug it into a project, first run it on a small task you won't mind losing for a week, and see its judgment error rate before deciding. Xiao He says she usually just wants speed, and even a few extra seconds of AI thinking feels slow; this "let it make the call on small things, bring in the big model for big things" sounds on point. If she really tries it, she'll start with something unimportant like writing a shopping list, not let it make decisions for her right off the bat.

Leaderboard Brief

All three boards haven't put out a new version this issue, with data frozen at October 2. The top spot on the human blind-vote text board is still Google's Gemini 4 Argon, but it has only racked up under 5,000 votes and its position is still shaky; the second place has over 70,000 votes and is far more reliable. The agent hands-on board tests "did it actually get the work done for someone," and the top spot is Anthropic's Fable 5.1. The image board's top spot is Anthropic's Fable 5, closely followed by Alibaba's Tongyi Qianwen 3.8, with the top five scores packed tightly — when picking tools, first check whether the number of rounds in parentheses is thick enough.

Section Picks (10 items)

  • A programming tool team says their idea is "everything is a plugin": it can be freely extended with plugins, easy to use and easy to mess up — before installing, first check who wrote what you're installing.
  • A 26-year-old piece of software was rewritten with new code by AI: when handing old software to AI for a refresh, first check whether you can verify every single change it made, one by one.
  • An exhibition whose title literally says "co-curated by AI": putting AI's involvement in the title is both honesty and controversy — looking at what it actually changed is more real than looking at slogans.
  • A practice partner that uses AI to simulate a German oral exam: using AI as a practice partner is nothing to be ashamed of, but don't take its oral scores too seriously — the real exam still depends on people and the real environment.
  • Masayoshi Son rarely issues a warning, and even he's uneasy about AI safety: when the people writing the checks start raising risks, it's worth a listen more than slogans shouted on stage.
  • AI assistants now have their own playground: treat new gimmicks as a spectator sport, don't rush to throw your own assistant in there.
  • Let your AI assistant cut a explainer video on its own: making videos is easier, but videos meant for the public still need your own once-over before posting.
  • A long piece asks whether AI really has "consciousness": talking like a human and truly understanding are two different things.
  • An analysis compares the US and China on "military AI," one loosening and one tightening: once the military door opens a crack, rule negotiations get tighter — anyone doing export business needs to keep an eye on it.
  • Someone wrote a piece called "AI makes me happy," pushing back against the wall of doom-and-gloom: the same tool can mean very different things to different people, and hearing one more voice doesn't hurt.

What Everyone's Watching

Scoring common tools to see which AIs can pick them up on their own; an AI assistant that sits beside you and takes notes during meetings; Chinese teams exploring letting AI think less and small models make the call; a programming tool branded "everything is a plugin"; an exhibition whose title says "co-curated by AI."

Tomorrow's Watch

① That tool scoring table — watch whether a third party re-runs the same batch of tools.

② The "large models think less, small models make the call" approach — wait until someone calculates the time and money saved before looking at it.

③ The programming tool branded "everything is a plugin" — wait until someone installs a pile of plugins and runs it for a while before saying whether it's good.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts