
LLMs can do work, but don't overhype general AI capabilities
I spent the weekend messing around with the topic "LLMs are real, AI is fake" and fell into quite a few traps. It started with that Pluralistic article on Hacker News. The title was provocative, claiming large language models are real, but so-called AI is fake. I treated it as a definitional argument and tested it with a few models at hand. The result? Product people can really get burned by this.
Let's clarify the terms first. LLM stands for Large Language Model—it's essentially a text machine that continues sentences based on probability. AI stands for Artificial Intelligence—the umbrella term is bigger, including search, recommendations, rule-based systems, and also the stuff vendors use to hype up general intelligence. There are two camps in the forums:
Some say LLMs are just one implementation of AI; others say so-called AI is just a prediction machine wearing a robot suit.
I ran two approaches. Approach A: Ask the model directly, "Is an LLM real AI? Is AI fake? Write me an explanation suitable for a WeChat Official Account." Approach B: Require it to break down concepts first, provide verifiable sources, distinguish facts from opinions and marketing fluff, and finally give a short conclusion.
For Approach A, I mainly used OpenAI. I've been using it daily for about two months and am pretty familiar with Q&A. I also tried a version with Daybreak Blue for three weeks; its structure for long Chinese texts was decent. The interface is just a chat box—paste the question, hit send. Responses were fast; out of my dozen or so questions, most took between 30 seconds and a minute. The problem is it loves to round off its statements. For example, it might write, "LLMs exhibit intelligence in some sense," which sounds philosophical but isn't verifiable. Two other responses claimed that "being able to call tools based on context" equals "being able to plan autonomously," which is a stretch. It just follows the most common human rhetoric.
The close-up of SelectionPlaneDriverPanel in my screenshots looks like a parameter panel, with temperature, max length, and web access toggles all laid out. Users might mistakenly think tweaking parameters turns the model from a text machine into a reliable agent.
Tweaking parameters doesn't change the knowledge boundary; it only changes the output style.
Approach B was much slower. I changed the prompt to: "Don't give pretty conclusions. First list source types, then label which come from papers, media, or forum discussions." The model outputs something more like a report, but source quality varies. It cited Reddit posts, Medium articles, and arXiv paper abstracts. I manually checked a few; the paper parts were solid, but the forum parts could only serve as opinion samples. Moonshot AI was something I tried for the first time in these three days; its Chinese expression is smooth and points are clear, but it also automatically adds phrases like "from a technical philosophy perspective." I've used DeepSeek for a month; it organizes complex statements well, but if you don't push it, the conclusions remain too rounded. I wrote about WorkBuddy before: don't start with research reports; start with a post-market daily report where sources can be checked. Looking back, that principle still holds.
In practice, Approach A takes about a minute, while Approach B takes two to four minutes because it needs to list sources and deconstruct opinions. I tested about 15 cases. In Approach A, two overstated the model's capabilities. In Approach B, one got the date of a news headline wrong, and another had sources that were too vague to click through. The gap isn't huge, but the direction is clear: without verification, LLMs are great at comforting you; with verification, LLMs start to become useful. So-called hallucinations happen when it prioritizes what sounds plausible over what is actually true.
So my conclusion is: it depends. If you just want a quick draft to edit, go with Approach A—it saves time. For science communication, reviews, or product docs, Approach B is safer. Don't treat LLMs as judges, and don't let them decide whether AGI exists. They are good as researchers, rewriters, or first-draft generators. The surprise with Approach B is that it separates pro and con arguments, but I still have to judge credibility myself. An LLM isn't an omniscient AI, but it can be a very competent text retrieval clerk.
Is it worth the price? With free quotas, Approach A is worth it because it's fast. Approach B is also worth it, but the value lies in reducing rework time, not making decisions for you. I'll mention the bad parts too: many products package LLMs as agents, pull a progress bar, write copy saying "planning autonomously," and users easily assume it's thinking. Often, it's just several prompts queuing up. The key is whether they expose sources, failure reasons, and manual verification entry points.
Going forward, I'll make the verification button a default option rather than something users have to click themselves.
📌 This article is compiled from Hacker News. Original: https://pluralistic.net/2026-09-12/god-in-the-box/
Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.
Physix Frontier