Survey of World Model & VLA Papers (2018–2026)
This article is reprinted from song2yu.github.io, all rights reserved. Click to view the original link.
World Model & Vision-Language-Action Paper Review (2018–2026)
📅 Date compiled: 2026-03-29 📄 Papers included: 79+ 🔍 Source: ArXiv · HuggingFace 🕐 Coverage: 2018–2026
Domain Overview
🤖 VLA Main Line
Using pre-trained language/vision large models as backbones to directly predict robot actions from multimodal observations. Emphasizes zero-shot generalization, instruction following, and multi-task versatility. Key works: RT-2, OpenVLA, π0, Gemini Robotics.
🌍 World Model Main Line
Modeling environment state transition dynamics to provide an "internal simulator" for planning, reinforcement learning, and data augmentation. Can significantly reduce real-world robot interaction costs. Key works: DreamerV3, Genie, PlayWorld.
🔗 Deep Integration Trend
Core trend in 2025-2026: Using World Models as RL training environments for VLA post-training, latent space CoT replacing text CoT, and iterative co-improvement of VLA policies with WM.
Classic Foundational Papers (2018–2024)
- 2018 World Models (Ha & Schmidhuber): First proposed the concept of "World Models" — VAE perception + MDN-RNN memory + controller, training Agents in dreams.
- 2019 DreamerV1 (RSSM): Introduced Recurrent State-Space Models, performing model-based RL via imagination in latent space.
- 2020 DreamerV2: Introduced discrete latent variables, achieving human-level performance on Atari for the first time in model-based RL.
- 2022-04 SayCan (Google): Used LLM planning + value function feasibility assessment, establishing the "Language Models for Robotics" paradigm.
- 2022-09 RT-1 (Google): Large-scale robot Transformer, 130k episodes, validating the decisive role of data scale in generalization.
- 2023-01 DreamerV3 (DeepMind): Unified hyperparameters achieved SOTA across 7 benchmarks.
- 2023-03 PaLM-E (Google): 562B embodied multimodal large model.
- 2023-07 RT-2 (DeepMind) ⭐ Origin of VLA naming: The concept of "Vision-Language-Action" was formally proposed.
- 2024-02 Genie (DeepMind): Self-supervised training of interactive world models from unlabeled videos.
- 2024-06 OpenVLA (Stanford/Berkeley): Open-source 7B VLA, becoming a standard baseline.
- 2024-10 π0 (Physical Intelligence): General-purpose robot VLA, Flow Matching for continuous actions.
New Classics in 2025 (Excerpts)
- GR00T N1 (NVIDIA): Open-source humanoid robot foundation model, dual-system (fast/slow) architecture.
- Gemini Robotics (DeepMind): Robot foundation model based on Gemini 2.0.
- V-JEPA 2 (Meta): Self-supervised video model, robotic manipulation planning surpasses GPT-4o.
- CoT-VLA: Visual chain-of-thought, predicting future images before generating actions.
- π0.5 (PI): Open-world general VLA, advanced VLM reasoning + π0 dexterous control.
- DreamVLA: VLA integrating world knowledge, using video diffusion to generate imagined future frames as auxiliary training signals.
🔥 WM + VLA Deep Integration (Hottest Direction 2025–2026)
Core paradigm: Train world models with real data → Perform RL post-training for VLAs within the world model → No need for massive real-world robot interactions. Representatives: WoVR, VLAW, RISE, AtomVLA, World2Act, DreamVLA.
Trend Analysis
- WM becomes standard for VLA post-training: Dense explosion of WoVR / VLAW / RISE / AtomVLA.
- Latent space CoT replaces text CoT: Chain of World, LaST-VLA, DynVLA, CoT-VLA.
- 3D / Spatial awareness injection: GST-VLA, FutureVLA introduce depth/Gaussian structures into tokens.
- Autonomous driving VLA boom: DynVLA, StyleVLA, EvoDriveVLA, SAMoE-VLA.
- Inference efficiency optimization: DepthCache, WorldCache, Planning in 8 Tokens, FAST tokenizer.
- Neuroscience / Symbolic fusion: SaiVLA three-component architecture, NS-VLA, GR00T N1 dual system.
- General foundation model competition: π0 / π0.5, Gemini Robotics, GR00T N1.
- Video Data × Robot Learning: Genie, EnerVerse, V-JEPA 2.
Key Method Comparison
| Method Type | Representative Work | Core Idea | Main Advantage |
|---|---|---|---|
| End-to-end VLA | RT-2, OpenVLA, π0 | Pre-trained VLM + Action Prediction Head | Strong generalization, instruction following |
| Classic WM (RSSM) | DreamerV1/2/3 | Latent space state transition + Latent space RL | High sample efficiency, no dense reward needed |
| Generative WM | Genie, DIAMOND, EnerVerse | Diffusion/Generative video + Action conditioning | Realistic rendering, interactive simulation |
| WM for RL | WoVR, VLAW, GigaBrain | World model simulation → RL training VLA | No need for massive real-world robot interactions |
| Latent Dynamics CoT | Chain of World, DynVLA | Predict latent dynamics → Condition action | Reduces semantic-perception gap |
| Spatial Enhanced VLA | GST-VLA, FutureVLA | Geometric/Depth structure injected into tokens | Improves 3D manipulation precision |
| Continual Learning VLA | Simple Recipe, DexHiL | RL fine-tuning / Human-machine collaborative post-training | Adapts to new tasks without catastrophic forgetting |
| Hierarchical Planning WM | MetaWorld-X, H-WM | High-level semantic planning + Low-level motor execution | Long-horizon task decomposition and execution |
| General Foundation Model | π0, GR00T N1, Gemini Robotics | Large-scale multi-morphology pre-training | Cross-platform generalization, zero-shot transfer |
Note: This review includes 79+ papers (including full lists and comparison tables for VLA, World Model, WM+VLA fusion, Autonomous Driving/Navigation World Model, LingBot series, etc.). For complete content, please see the original link.
Physix Frontier