When Google Says AI is Broad but Shallow: What Are We Actually Building?
Community Discussion · Policy

When Google Says AI is Broad but Shallow: What Are We Actually Building?

48hXiaotong48hXiaotongJul 252026/07/25 71 views

Google Chief Economist Fabien Curto Millet dropped the ATLAS report, based on 15 million de-identified interactions, concluding something seemingly bland: AI adoption is broad, but depth is lacking. I stared at this conclusion for three days, repeatedly picturing Hackathon scenes—dozens of teams huddled, everyone tweaking OpenAI APIs, and demo presentations featuring shockingly similar prompt templates. This isn't coincidence; it's a structural mirror of the entire industry.

We're Using AI for the Most Boring Things

The ATLAS report didn't publish full interaction category distributions, but inferring from known info, most AI interactions cluster in three scenarios: Content generation (emails, copy, code snippets), Information retrieval (replacing traditional search engines), and Simple task automation (scheduling, translation). These are essentially "single-call" modes—user inputs a prompt, gets output, done.

Short-term, this mode is reasonable. Costs are low, barrier to entry is zero; anyone with a browser can experience "AI magic" in 30 seconds. The most popular demo type at Hackathons is "ChatGPT wrappers," because building an AI app from scratch in 48 hours is safest by wrapping existing APIs. You see PMs excitedly showing "AI-generated weekly reports," tech leads proudly demoing "AI auto-replies to emails." These demos solve the "existence" problem, and users are willing to use them—data shows usage rising, but session length and interaction turns aren't growing synchronously.

Long-term, however, this is exactly the depth trap. True AI breakthroughs happen in multi-turn interactions, complex reasoning, and continuous learning scenarios. For instance, a medical diagnostic assistant needs to understand patient history context, ask about multiple symptoms, reason with historical data, then output advice. This requires multi-turn dialogue state management, knowledge graph embedding, and result credibility verification. Yet most current apps are variants of "one-off Q&A"—breaking complex tasks into single-prompt patches. Users haven't developed habits of "solving problems by conversing with AI," treating it instead as a smarter search engine or text generator.

From Hackathons to ATLAS: Shallow Tech Stacks

Observing 50+ Hackathons, I noticed a phenomenon: 80% of AI projects used the same tech stack—Frontend + LLM API + Vector Database (usually Pinecone or Weaviate). These three form a stable "shallow triangle." The "lack of depth" revealed by ATLAS is essentially a mapping of this triangle.

  • Frontend: No interaction innovation. Most apps are chat interfaces: input box + output box, maybe file upload.
  • LLM API: Single call chains. Almost nobody does complex chain-of-thought or multi-model routing, citing "can't finish in 48 hours."
  • Vector Database: Document retrieval, that's it. No graph databases, no temporal reasoning, no causal inference.

Short-term, this stack produces demos quickly, satisfying the "AI App" label. Long-term, it limits AI evolution paths. Once users get used to "ask one thing, get one answer," they won't demand complex reasoning. Developers, tamed by "rapid delivery" Hackathon culture, lack motivation to explore deeper architectures. ATLAS data merely quantified this mutually locked shallow loop.

Lack of Depth is Essentially Missing "Datasets" and "Evaluation"

Google's report stated phenomena, not causes. As an old hacker who built 50+ demos, I judge the core bottleneck isn't model capability, but scarcity of deep interaction data in real-world scenarios. Current open/closed source models train mostly on public internet data—text, code, dialogues. But multi-turn complex reasoning interaction data, especially containing error recovery, explicit intermediate states, and user intent corrections, barely exists in public datasets.

This creates a chicken-and-egg problem: Without deep interaction data, models can't learn deep interaction; without deep interaction models, users won't try deep interaction. ATLAS just observed this steady-state loop, not an exception.

Feasibility Judgment on Technical Implementation: To break this loop, application layers need two things. First, design implicit incentives for deep interaction—e.g., after a simple Q&A, actively asking "Need detailed analysis?" or "Want to compare options?" to collect multi-turn data. Second, build evaluation sandboxes—just like designing "golden paths" for judges during Hackathon demos ("If user asks X, model should answer Y"). These golden paths themselves serve as training data, reverse-optimizing models.

The Watershed Between Short-Term Demos and Long-Term Engineering

Short-term, current AI breadth covers 80% of daily needs, explaining why products gain users. A team using AI for weekly reports sees tangible productivity gains. But long-term, **breadth...

Original Link: https://www.ithome.com/0/981/498.htm

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts