The physical video generation variant of the Concept World Model. It improves upon state-of-the-art benchmarks
A world model learning objects, states, actions, and causality in latent space. Unlike pixel-generating models