When AI Writes Poetry But Can't Hold a Cup: Physical World Models are the True 'Hardcore' Track
Have you ever thought about this? AI can write poetry, paint pictures, and help you draft weekly reports. But if you ask it to control a real robot to walk to a table, pick up a water cup, and hand it to you, it might crush the cup, knock over the table, or simply fail to find where the cup is. This scenario isn't from a sci-fi movie; it's currently AI's most awkward weakness—it doesn't understand physics.
In 2026, world models have become the most crowded track in the AI circle. But most players are still just slapping a shell on Large Language Models (LLMs) or using video generation models to pretend they understand physics. It wasn't until Zhang Lihua, the founder of NVIDIA PhysX, entered the game and announced his self-developed next-generation physical world model, Fysiverse, that I felt someone who truly understands physics has finally come to draw the line in this track.
Why is "Physical AI" the Achilles' heel for robots?
Let me start with an example I tested personally. Last year, I took a demo of a so-called "strongest robot manipulation model" and asked it to command a robotic arm to grab an egg. The result: the model recognized the label "egg" and directly output a "grab" action—the robotic arm squeezed down hard, and yolk flowed all over the table. The reason is simple: the model only knows "egg" is an object, but it doesn't know eggs are fragile, nor does it know how much force will cause them to break.
This is the fatal flaw of current AI: lack of physical common sense. LLMs learn text associations, and video generation models learn pixel distributions. Neither can learn physical laws like "gravity," "friction," or "elastic deformation." For robots to operate in the real world, they must understand how objects behave under force.
What Zhang Lihua is doing with Fysiverse is essentially building a hybrid system of "physics engine + AI." Unlike traditional game engines that only simulate rigid body collisions, it uses neural networks to learn the physical properties of materials, such as the softness of cloth, the fluidity of water, and the compressive deformation of sponges. Once this type of model matures, robots can "rehearse" actions in virtual environments before deploying them in reality, eliminating the need for brute-force trial and error.
From a propagation perspective, this topic is practically a traffic magnet for short videos. Viewers are already aesthetically fatigued by "AI writing poetry," but contrasting scenes like "robots crushing cups," combined with the hardcore concept of "Physical AI," easily spark discussion. I'm planning to make a comparison video: on the left, show realistic images of "a water cup on a table" generated by AI; on the right, show the actual footage of a robot breaking the cup when trying to pick it up. Finally, introduce Zhang Lihua's solution. Isn't that more interesting than "AI learned another new joke"?
From PhysX to Fysiverse: A Veteran Physics Engine Expert's Dimensional Strike
Who is Zhang Lihua? He was one of the founders of NVIDIA PhysX. What is PhysX? It's the physics engine used in almost every AAA title (Assassin's Creed, The Witcher 3). Falling rocks, fluttering cloth, and exploding debris—all simulated by PhysX. He accumulated over 20 years of experience in physics simulation within the gaming industry. Now, transplanting this experience into the AI field is essentially a "dimensional strike."
But note, there is a fundamental difference between game physics engines and physical world models. Game engines pursue "visual effects"—as long as the cloth looks good while fluttering, there's no need to precisely simulate the stress on every fiber. But for a robot to grab cloth, it must know where to grip without slipping and how much force to apply. So Fysiverse isn't just a simple "physics engine + AI"; it uses data-driven methods to let the model learn physical laws.
I'm particularly interested in his technical route: not end-to-end learning with pure neural networks, but first building a "differentiable" simulation environment with a physics engine, then integrating neural networks into this environment for training. The benefit of this method is that the model won't hallucinate (e.g., making water flow uphill), because the underlying physics engine limits the possibilities. The downside is that building such a simulation environment requires massive computational power and a large amount of real-world physical data for calibration.
This leads me to a more fundamental question: Should physical world models rely on "brute force miracles" like GPT, or "manual modeling" like traditional simulations? Zhang Lihua chose the middle path, but can this road work? If it succeeds, it may have greater commercial value than simple video generation models—because robot manufacturers, autonomous driving companies, and manufacturing industries all need such a "physical common sense training ground."
Traffic Perspective: Why Will Audiences Pay for "Physical AI"?
I estimate the propagation potential of this news lies in "cognitive contrast." Audiences are accustomed to the narrative that "AI is omnipotent," but few know that AI is an "idiot" in the physical world. You can use this contrast to create content:
- Opening: Show a video of AI writing poetry and painting, then suddenly cut to a clip of a robot smashing a vase. Maximize the contrast.
- Middle: Explain what "Physical AI" is using accessible metaphors, e.g., "AI learning physics is like you learning swimming; watching endless tutorials without getting in the water is useless."
- Ending: Predict how Fysiverse will change the robotics industry, while casually roasting many current "pseudo-world model" projects.
However, be careful. This kind of hardcore tech content easily falls into the trap of "bombarding viewers with technical jargon." Audiences don't understand what a "differentiable physics simulator" is, but they understand "letting the robot break 1 million virtual cups first, so it won't break a real cup when gripping it." The core propagation logic is: translate professional terms into "common sense + stories."
Additionally, Zhang Lihua's background itself brings traffic. The label "NVIDIA PhysX Founder" has recognition among both gamers and AI enthusiasts. If he opens an account on Douyin or Bilibili to do science popularization, I think he would be more persuasive than many AI bloggers.
Actionable Advice for Readers
If you are a tech short-video creator, I suggest grabbing this topic immediately. Don't be a "news parrot," be a "cognitive disruptor." Find a specific physical action nearby (like pouring water, folding clothes, twisting a bottle cap), record a failure case of a robot performing this action with your phone (if you don't have one, find public compilations of robot fails), and then compare it with Fysiverse's solution. Audiences love the feeling of sudden realization: "Oh, so that's where the problem was."
If you are an ordinary viewer, next time you encounter any AI product claiming "it can teach robots everything," ask yourself first: Does it understand physics? If not, it's likely just riding the hype.
Spring hasn't arrived for Physical AI yet, but Zhang Lihua's entry at least gives this winter a reliable...
Physix Frontier