Community Discussion · Tracks

ReviewRadar · 2026-10-03

Issue 066 is here. Three highlights this issue: one about how the top-ranked model on public leaderboards isn't necessarily the best fit once you bring it into real work, one about AI helping a database dig up old vulnerabilities that had been missed, and one about NVIDIA putting out a small box you can put on your desk that can run large models locally.

Key Updates

1. The top-ranked model on the leaderboard isn't necessarily the best fit once you bring it into real work

An evaluation company wrote a hands-on piece about how picking an AI can't just rely on public leaderboard scores. Leaderboards test with a uniform set of questions, but your work is long and messy. The same question it aces might, when swapped for the pile of stuff you actually need to handle, do worse than the fifth-place model on the board. Old Xu's take: public leaderboards are like a standardized exam, but the work in your project has fuzzy boundaries and weird errors that the exam can't cover. When he picks coding tools, he always starts by taking a small module he wouldn't mind losing and rolling with it for a week, running the same batch of work through it and his current setup, and whoever makes fewer mistakes wins. Xiao He says from now on, whenever someone tells her some AI ranks first, she'll first try it on her own everyday little tasks, and only switch if it works well.

2. AI helps the database dig up old vulnerabilities that had been missed

MariaDB is a commonly used open-source database, and many websites store their accounts and orders in this kind of database. This year they used AI to help scan the code for security flaws, and it turned up quite a few vulnerabilities that humans had looked at for a long time without finding. Old Zhou reminds us that AI finding vulnerabilities is just an extra pair of eyes, it can't replace human judgment. What it digs up needs human review, and whether it can actually be exploited and how big the impact is still has to be decided by a person. If you use it to scan your own database, don't start by touching production data—run it on a read-only copy first. Xiao He says she used to just wait for official patch notices, and now she knows there's this proactive vulnerability-scanning use case, but for important systems she'll still wait for a dedicated person to confirm before acting.

3. NVIDIA put out a small box you can put on your desk that can run large models locally

NVIDIA released a small machine called DGX Spark, with 64GB of memory, aimed at "running AI locally." If you want to use a decent open-source model, you no longer need to rent that distant machine in the cloud—this thing on your desk can do it, and your data never leaves your home. A-Kai says what he watches for is never the performance number on the official site, but the three things beyond the spec sheet: does it get hot at full load, is the fan noisy, and how long can it last on a charge. 64GB can run quite a bit more than an ordinary laptop, but it's ultimately a small host machine—before buying, think clearly about whether you'll run it every day or get bored after two days and let it gather dust. Xiao He says the line "your data never leaves your home" is useful to her, and she'll consider it once the price drops and someone writes up the heat and electricity costs.

Leaderboard Flash

This issue's main board has switched to the Artificial Analysis model intelligence index leaderboard, already updated. The top spot is Anthropic's Claude Opus 5.5 with 58 points; OpenAI's GPT-6 Astra and Google's Gemini 4 Argon are both in the 53-point tier; the highest among open-source is Xiaomi's MiMo-V2.6-Pro with 46 points. Also, the two old Arena boards (human blind-vote text board, coding board) weren't updated this issue, with data as of 2026-09-30 and 2026-10-01 respectively, so don't use the rankings to pick tools yet. When picking a coding AI, first check whether the number of matches in parentheses is enough.

Section Picks (10 items)

  • An AI assistant sent your entire key bundle out: before letting AI touch files, think clearly about what absolutely must not go into its packing list.
  • Another open-source "fast report writing" model: speed is its selling point, but whether the report is any good still needs you to try it with your own materials.
  • One interface that picks which model you should use for you: auto-picking models saves effort, but you may not know who's ultimately answering your question.
  • Let AI write tests for you, and it ends up testing nothing: tests aren't better the more you write, only the few that can catch errors count.
  • Give AI a brain that "understands your company": making AI understand your company's rules is more practical than swapping in a smarter model.
  • In the client room, AI revises a version first, and it only counts once a person nods: AI proposes, people decide—this division of labor is less likely to go wrong.
  • Give it one sentence, and AI builds you a LEGO model: this kind of build-along-with-you tool, the fun is in the process, not in getting it done in one step.
  • Someone wants to add a "protect the kids" gate to AI: the big risks haven't arrived yet, and the small harms right in front of you are what most families hit first.
  • Two photos, generate a "hotel lobby rap": the easier generation gets, the more you need to think clearly—using someone else's face to make a video requires consent first.
  • Build a "behind closed doors" AI workbench for the team: self-hosted, with clear permissions, is more reassuring than having more features.

Everyone's Watching

The top-ranked model on the leaderboard isn't necessarily good once brought into real work; AI helps the database dig up old vulnerabilities that had been missed; someone's AI assistant sent the entire key bundle out; NVIDIA's local AI small box, 64GB, sits on the desk; give AI a brain that understands your own company, and someone wrote up the whole implementation process.

Tomorrow's Watch

① For the piece about the top-ranked not necessarily being the most usable, watch whether anyone runs a comparison again with Chinese models.

② For the batch of vulnerabilities MariaDB dug up with AI, wait for the follow-up patches to land and see if anyone reproduces them.

③ For NVIDIA's local AI small box, wait for people who get their hands on it to write up the heat, noise, and electricity costs.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts