Do World Models Need an Explicit Physics Engine?
I've recently read several papers on video-based world models, such as Hobbes, Genie, and follow-up analyses of Sora. One core question has been bothering me: Is the "physical intuition" learned by these models truly causal understanding, or just statistical correlations fitted onto visual manifolds? Take scenarios like objects falling or colliding; while purely data-driven diffusion models can generate nice-looking trajectories, they often break down when encountering rare cases in the training set, like balancing long rods or non-rigid deformations.
On the other hand, approaches that explicitly embed physical constraints, like Neural-ODEs or Hamiltonian networks, generalize well but are limited by predefined equation forms, making it hard to cover complex scenarios involving friction or fluids. I lean towards a compromise route—using differentiable physics engines as the "underlying operators" of the world model, with neural networks learning residuals or unmodeled effects at the upper layer. Do you think this hybrid architecture would be constrained by computational real-time requirements in actual deployment (e.g., robot planning)? Or are there better ways to inject priors?
Physix Frontier