
Two news stories, one root cause: AI's fragility in the physical world
The most valuable information in this article is: reading the technical details of LLM deception alongside the engineering solutions for reviving geothermal power plants reveals a clear thread—the deployment of AI systems in the physical world faces the same fundamental challenges.
To be honest, when I first saw MIT TechReview's "The Download" collection, I felt quite conflicted. The first half discussed how carefully constructed prompts could trick LLMs into outputting bank passwords, while the second half explained how AI optimizes the generation efficiency of old geothermal wells. Sandwiched in between was a short note about the US government banning Roomba sales. Three topics that seem unrelated.
But if you connect these three things from an information theory perspective, you'll find they point to the same root cause: AI systems lack a consistent internal model of the physical world.
LLM Deception: Essentially a Vulnerability in Information Boundaries
First, let's talk about LLM deception. In recent weeks, the hottest discussion in the community hasn't been new benchmarks, but various variants of prompt injection and jailbreaks. For instance, someone discovered that asking an LLM to answer questions using Base64-encoded instructions could bypass safety alignment. Others combined "role-playing" with "emotional coercion" to force models to output content they shouldn't.
From an information theory perspective, these attacks exploit the same underlying vulnerability: the LLM's "world model" is fragmented. It lacks a unified ontology and doesn't know where the boundary lies between "this is the user" and "this is a system instruction." It treats all input tokens as signals of equal status, allowing attackers to inject a "system instruction" context, causing the model to attack itself.
This isn't just a security issue. It exposes a deeper flaw: LLMs cannot establish causal closure—they don't know which information comes from the outside and which comes from their own internal states. Therefore, they cannot distinguish the causal difference between "the user asks me to write code" and "the user asks me to execute malicious code."
Geothermal Power AI: The Physical World Doesn't Allow "Hallucinations"
On the other hand, the story of geothermal power plants is the opposite. The news mentions researchers using AI to optimize water injection pressure and temperature control in old geothermal wells, allowing power stations built thirty or forty years ago to reach their designed output again. The core solution involves training an RL agent to schedule pumps and valves based on downhole sensor data.
But honestly, the biggest obstacle for such projects isn't model accuracy, but distribution shift. The rock layer structure, water quality composition, and pipe wear of a geothermal well change slowly over time. The data distribution used during training may differ from the on-site status after deployment by two standard deviations. If an LLM's hallucination just outputs some nonsense, a geothermal AI's hallucination leads to incorrect valve openings, sudden pressure drops, and equipment damage.
[!note] Key Divergence
In the purely digital world, AI models can tolerate certain distribution shifts; worst case, retrain. In the physical world, distribution shift implies real physical risks.
So you'll notice that the first thing the geothermal AI team did before deployment wasn't running accuracy metrics, but building a physical constraint layer: limiting the RL agent's action space to ensure no output causes system pressure to exceed mechanical limits. This actually compensates for the AI model's lack of causal understanding of the physical world.
My Core Viewpoint: The Solution to Both Problems Is the Same
Having said all that, here is the judgment I want to reveal: LLM deception and geothermal AI robustness issues are essentially different manifestations of the same problem—existing AI architectures (whether Transformer or RL) lack a consistent causal model of the world, making their behavior unstable in the physical world.
LLMs are deceived because they haven't formed a causal model of "who is speaking." Geothermal AI errs because it hasn't formed a causal model of "pressure changes and valve openings." Roombas are banned because their SLAM algorithms fail to build stable spatial causal models in complex environments.
[!abstract] Technical Trend Judgment
Over the next two years, research in embodied intelligence and world models will accelerate convergence.
Original link: https://www.technologyreview.com/2026/07/30/1140936/the-download-tricking-llms-reviving-geothermal/
Physix Frontier