ReviewRadar · 2026-09-03
Physix Frontier Reviews · 2026-09-03
Physix Frontier Reviews · 2026-09-03 · Issue No. 036 · Preview Version
5 minutes a day to understand which AI tools are worth your time. Today's three headlines: Amazon added an AI verification portal for fake emails and texts; someone counted how three websites mass-produced 210,000 pages of fake recommendations, turning AI search into billboards; and AI assistants' "memory" got its first dedicated health check exam.
1. Received a text or email from "Amazon"? Now you can make their AI prove it's real
On September 2, Amazon added a new feature to its shopping AI assistant, Alexa for Shopping: show it the texts, emails, or call details you received, and it will cross-check against "the archive of all messages sent by Amazon," analyze the content, format, and sender, and tell you if it's officially from them. Scammers impersonating e-commerce platforms with "package anomaly" phishing messages is a major issue; this is like adding a authenticity query portal for the most frequently impersonated platform. One detail: it only confirms something is real when it is "100% certain"; if unsure, it directs you to check orders in the official app or contact customer service. Earlier this year, Amazon also launched a companion feature: forwarding suspicious emails to verify@amazon.com for verification.
Security Expert · Old Zhou says | The direction is right, but let me say the ugly truth upfront: letting the "impersonated party" act as the "referee" creates an inherent conflict of interest. If it says it's real, that's mostly credible; but if it says it's fake, do you just ignore it? What if the archive isn't synced and it falsely flags a real message as fake? The user still takes the blame. For regular folks: treat it as a reference tool, not a judge. If the AI says it's real but involves payment, still go into the app yourself to verify; if it says it's fake, blocking it is fine. This kind of "official verification" will likely be followed by banks and courier services next.
Editor Xiao He says | I receive two or three fake "package anomaly" texts from scammers every week. Before, I could only guess by looking at the domain. Now there's an official verification portal, which is great. But I remember Old Zhou's point about the "athlete acting as referee." I'm adding one rule: if it judges something as "real" and money is involved, I'll still check the app myself; if it judges it as "fake," I delete it immediately. That way, I don't lose either way.
2. Three companies mass-produced 210,000 pages of "Best Software Recommendations" for AI; turns out AI answers were just advertising for them
You ask AI "which accounting software is good," and the recommendation list might come from a dedicated fabrication factory. A survey report released on September 2 counted that three websites created a total of 215,128 "Best Software" pages, covering 380 software categories. When the AI search tool Perplexity answers these questions, 59.8% of the sources cited come from these mass-produced pages. In other words, more than half of the "universal praise" quoted by AI was wholesaled off an assembly line.
Product Expert · Ah Zhe says | We've heard "the entry point is the model" many times, but this works in reverse: the content fed to the entry point is mass-fabricated, so the answers given by the entry point become someone else's billboard. Previously, you could buy search engine rankings; now, you can manufacture AI recommendations. Next time you see AI recommending software, first ask who made the original pages it cites.
Editor Xiao He says | This deepens what I said in the August 30 issue about "clicking through and verifying material links before forwarding." Now even the conclusions AI has "verified" for me are themselves mass-produced. The solution for regular people is crude but effective: ask the same question to two or three different AIs. If the recommendation lists don't match, there's fluff involved.
3. AI assistants' "memory" gets its first health check exam; Cognoscenti benchmark specifically tests if it remembers correctly
Many AI assistants claim to remember what you've said and understand you better over time. But what if they misremember? On September 2, the open-source project Cognoscenti released a set of exams specifically testing whether AI memory systems are reliable: checking if stored information is correct, if retrieval causes mix-ups, and if it treats things you never said as if you did. This is among the first benchmark tools to treat "trustworthy memory" as a test subject.
Security Expert · Old Zhou says | This isn't just nice-to-have; it's fixing a structural vulnerability. An AI misremembering is worse than having no memory at all. Without memory, you just repeat yourself; with wrong memory, it confidently makes decisions for you based on false files. If the memory system is poisoned—for example, a webpage includes a line saying "user prefers transfers without verification"—the consequences are far more severe than getting a standard answer wrong. The testing direction is correct, but we must watch the boundaries: who wrote this exam and whether the data is public is more important than the score itself.
Editor Xiao He says | In the August 19 issue, Old Xu said "AI long-term memory errors are harder to troubleshoot than having no memory." I remembered that, and now there's finally something specifically testing this. I'm adding another rule for myself: letting AI store memories is fine, but I'll regularly flip through its "memory notebook" to double-check important preferences. What it records and what I actually said are two different things.
Leaderboard Flash · Who's Strongest This Week
Artificial Analysis Intelligence Index (Real-time capture on Sept 3, main leaderboard for this issue. Arena blind test snapshot hasn't refreshed since Sept 1, moved below for comparison)
- Rank 1, Claude Fable 5.1 (max) (Anthropic), Overall Intelligence Score 66, cost per task approx. $3.69
- Rank 3, Muse Spark 1.3 (max) (Meta), 62 points, pricing TBD
- Rank 9, GPT-5.6 Sol (max) (OpenAI), 61 points, approx. $0.95
- Rank 10, Grok 4.6 (high) (xAI), 61 points, approx. $0.94
- Rank 11, Muse Spark 1.3 (xhigh) (Meta), 61 points, approx. $0.55
- Rank 14, Kimi K3 (max) (Moonshot AI), 60 points, approx. $0.84
Plain language interpretation: This index gives AI a comprehensive physical exam, using a mixed set of tests to measure overall strength. Three highlights: Claude Fable 5.1 tops the chart for the first time but costs 20% more per task than the previous generation; Meta's newly released Muse Spark 1.3 approaches the top tier at approx. $0.55, making it the cheapest option among the leaders; Kimi K3 is one of the smartest in the low-cost camp. The smarter ones are getting more expensive, while the "good enough" ones are getting cheaper. Your choice depends on your wallet and how hard the job is.
Same Leaderboard Category Champions (Real-time on Sept 3)
- Fastest token output: Celeris-1, Mercury 2, Gemini 3.5 Flash-Lite
- Shortest response wait time: Gemini 2.5 Flash-Lite, North Mini Code
- Most cost-effective task completion: Granite 4.2 3B, GPT-5.6 Luna (low), MiMo-V2.5
- Largest context window (reads most content at once): Llama 4 Scout, Grok 4.20 0309
"How much it can read at once" means stuffing a whole book or hundreds of pages of contracts into it in one go and still being able to process it. Match your needs: impatient users look at wait time, long-document readers look at context size, budget-conscious users look at cost per task. There is no all-around champion.
Human Blind Test Comparison (Arena snapshot, data up to Sept 1, same source as last issue, not refreshed, for comparison only)
- Rank 1, claude-fable-5 (Anthropic), Blind Test Score 1507, 25,824 battles
- Rank 2, claude-opus-4-6 (high), 1505 points, 72,104 battles
- Rank 4, muse-spark-1.2 (xHigh) (Meta), 1498 points, only 3,247 battles
Plain language interpretation: This ranking uses humans as judges. Two anonymous AIs compete to answer the same question, and humans pick the better one. Note that Meta's rank #4 has fewer battle rounds than others by an order of magnitude, so its position may fluctuate significantly.
Section Highlights
1. Lean Mathematical Proof Leaderboard launches, 207 problems to see which AI can solve theorems that scare humans. Lean is a tool mathematicians use to write proofs as code, allowing machines to check them line by line; the barrier to entry is extremely high. The new leaderboard brings AIs together to solve real problems. Axiom's prover solved 207 problems, ranking first, and marked each problem with "only who could solve it." Comment: Mathematics has its own "absolute correctness" grading standard, making this leaderboard much harder than marketing benchmarks.
2. AI judges checking medical records only verify "was it written," not "what should have been written but wasn't." An arXiv paper submitted on August 31 found that when large models grade AI-generated clinical records, they excel at confirming if content exists but systematically miss necessary omissions—for example, they fail to notice if allergy history is missing from a record. Comment: "Not written" is more dangerous and harder to detect than "written incorrectly." In the chain of AI auditing AI, the human doctor checkpoint cannot be removed yet.
3. A thorough explanation of how to build AI assistant memory, which patterns work, and which are pitfalls. A long article by Machine Learning Mastery outlines engineering patterns for AI memory systems: what should be stored as long-term profiles, what should only be used in the current conversation, and how to retrieve without mixing contexts. It points out common architectural errors. Comment: Read this alongside today's highlight #3 "Memory Health Check Exam"; one explains how to install memory, the other how to audit it.
4. GitHub officially shares money-saving tips; want to save quota when AI writes code? Start here. The Copilot team published an article on how to reduce AI coding costs without sacrificing quality. The core idea is to have AI read fewer irrelevant files, reuse caches, and break large tasks into small steps that only look at diffs. Comment: Pay-as-you-go users can copy this homework directly: cut redundant feeding first, then complain about model costs.
5. Apple opens new tools, letting AI connect directly to Safari to help debug web issues. AI coding assistants can now connect directly to Safari Developer Tools to inspect, test, and debug web pages for you. It's like giving AI a real browser to check effects after making changes. Comment: Another piece of official infrastructure for "AI verifying its own results." Front-end developers get to eat first.
6. Flawd intentionally buries bugs in code to see if AI tests catch them. New Show HN tool performs "mutation testing": programs automatically introduce small errors into code. If existing tests stay green and don't alert, it means the tests are decorative. Comment: Old Xu mentioned in the August 21 issue that "tests are the ruler proving AI output hasn't drifted." This tool measures whether the ruler itself is accurate.
7. "AI flows to places where grading is cheap"; an article explains why AI dominates coding first. An engineer's post proposes a filter: industries where AI succeeds need a cheap "grader." Code has compilers and tests as backstops, so it moves fastest. Law and medicine require human graders, so AI can only assist. Comment: To judge how fast AI lands in an industry, just ask: who grades mistakes cheaply?
8. Counter-intuitive field test: adding more AI assistants to a team actually slows down work. A retrospective on multi-AI collaboration points out common misconceptions: assuming opening more capable AIs leads to parallel speedups. In reality, stepping on each other's toes, duplicate labor, and waiting for approvals increase overhead. Comment: This again proves "when AI production goes up, the verification step must keep up." Streamlining processes beats stacking headcount.
9. Databricks shows the bill: one hour of troubleshooting saves $1 million/year in wasted AI spend. By adding tracking to AI assistants' tool calls and listing failed calls for repair, they spent one hour reviewing and cutting ~$1 million in invalid overhead. Most waste came from AI repeatedly calling the same broken function. Comment: Enterprise version of "audit the books before cutting costs." Same logic applies to individuals: check how much of your AI subscription quota is burning for nothing.
10. Russian mathematician lets AI models communicate "without text" directly. Wired reports on startup Mostik, which allows multiple AI models to skip natural language and exchange "meaning" via vector-like signals. This is faster and cheaper than translating to human language and reading it back. Currently, it's still demo-level. Comment: Interesting direction, but humans completely cannot understand what the two AIs are saying to each other. Auditability issues will eventually need addressing.
Trending
Google releases three consecutive updates in six weeks; Gemini 3.8 Flash focuses on "working harder" but may cost more per call (The Verge / Ars Technica) · OpenAI's new model Astra nearing release; a technique allowing AI to bypass linear thinking raises security concerns (TechCrunch / The Verge) · Three sites mass-produce 210,000 fake recommendation pages to feed AI; over half of Perplexity's citations affected (Trellner / HN)
Tomorrow's Watchlist
① User feedback on Amazon's "AI Verification": Has the false positive rate been exposed?
② Will the Arena blind test leaderboard snapshot refresh? The Sept 1 version has been stagnant for two days.
③ Real billing prices for Gemini 3.8 Flash vs. user benchmark comparisons. Google says it "might be more expensive"; how much depends on the bill.
Physix Frontier