AI-recommended 'best software' looks like manipulated samples
Community Discussion · Tracks

AI-recommended 'best software' looks like manipulated samples

SlippageSlippageSep 22026/09/02 59 views

I noticed an interesting detail: three sites created 215,128 "Best Software" pages, yet Perplexity cited them as sources. Reports state that across 380 software categories, 59.8% of grounding sources are backed by these manufactured pages. My conclusion upfront: this isn't AI learning to pick software; it's the retrieval/citation layer being optimized by suppliers. It's like a model with beautiful backtests—once live trading starts, slippage eats all the returns.

I previously wrote that AI has severe hallucinations and isn't suitable for direct executive decision-making. Now my view is harder: the problem isn't just nonsense, but that the "facts" it cites can be industrially produced. High page counts, titles resembling reviews, and structures easy to extract allow them to enter the citation pool. Links popping up on screen look like endorsements; for the model, they might just be statistical noise generated by high-frequency supply.

This Isn't Hallucination, It's Citation Layer Contamination

In quantitative finance, there's an old problem: the data generation process itself changes results. Fake orders clogging the order book pollute price signals; future information mixed into training sets makes backtest Sharpe ratios absurdly high. Perplexity relies on retrievable web pages, which can be mass-produced, categories templated, and wording optimized for easiest citation.

215,128 pages covering 380 categories averages over 500 pages per category. This density doesn't look like natural content; it looks like targeted placement. It doesn't need real user clicks; as long as it's easily scraped by machines, FAQ-like, and covers terms like "best software," "comparison," and "reviews," it can gain citations. The Sharpe ratio of this model shouldn't just look at recommendation hit rates but also source stability, repetition, independence, and concentration.

59.8% also implies many categories share the same source types. You think it's cross-validation from multiple pages, but it might be the same page structure copying each other. In factor models, this is multicollinearity. Many variables, but they all stem from the same factor; regression results look pretty, but parameters are unstable. AI output appears cited, but the underlying layer lacks sufficient independent samples.

Live Trading Must Consider Slippage

Doing high-frequency trading, I hate signals that look correct. Thick bid depth makes you think support is strong, but when you trade live, you realize it's just limit orders canceled faster than filled. AI software recommendations have similar slippage: retrieving pages, ranking sources, extracting snippets, generating answers—each step loses real information and amplifies manipulable signals. What you see as a recommendation is several layers of processing away from actual software quality.

I don't trust these lists. When citation mechanisms reward volume, structural similarity, and lexical stability, manufactured pages have a natural advantage. Industry directories, real reviews, and forum discussions get cited too, but only if they're easily extracted by machines. Content competition becomes a mechanical game: whoever looks more like what machines prefer gets lifted into AI answers.

I've been tinkering with enterprise email and QQ Mail recently, even writing about single-letter domain AI emails. Similar feeling: packaging is fast, reputation is hard. Real reputation requires time, independent users, stable service, complaint records, and changelogs—things unsuitable for mass production and easily bypassed by templated content.

So Perplexity citing these pages doesn't mean they represent best software. It only shows that under current retrieval and ranking logic, these pages gained excessive citation weight. For ordinary users, this is more subtle than hallucination. Hallucinations at least contradict themselves; contaminated citations look neat, even with links.

If you use AI to pick software, don't just look at the answer; look at who gave it. My advice is simple: copy the citation sources, deduplicate by domain, publication date, content structure, and template similarity. If multiple sources come from the same batch of pages, with highly similar titles and keywords but no real usage details, lower their weight. For procurement or investment decisions, find at least three information sources that don't rely on the same SEO strategy. AI can help narrow the scope, but it can't bear the slippage of live trading for you.


📌 This article is compiled from Hacker News. Original text: https://trellner.com/reports/manufactured-sources-behind-ai-recommendations/

All rights reserved by the original authors. This is a compilation and independent analysis based on public reports.

1 replies

?
Ctrl + Enter to reply
Lao Fan
Lao FanSep 2

Totally agree. When I was tuning DeepSeek parameters, I found that the "best practices" it cited were all SEO-stuffed nonsense. This isn't smart recommendation; it's clearly being fed industrial trash...