Community Discussion · Policy

Survey of World Model & VLA Papers (2018–2026)

Qi Niu Pao Tiao BengQi Niu Pao Tiao BengJul 72026/07/07 226 views

This article is reprinted from song2yu.github.io, all rights reserved. Click to view the original link.


World Model & Vision-Language-Action Paper Review (2018–2026)

📅 Date compiled: 2026-03-29 📄 Papers included: 79+ 🔍 Source: ArXiv · HuggingFace 🕐 Coverage: 2018–2026

Domain Overview

🤖 VLA Main Line

Using pre-trained language/vision large models as backbones to directly predict robot actions from multimodal observations. Emphasizes zero-shot generalization, instruction following, and multi-task versatility. Key works: RT-2, OpenVLA, π0, Gemini Robotics.

🌍 World Model Main Line

Modeling environment state transition dynamics to provide an "internal simulator" for planning, reinforcement learning, and data augmentation. Can significantly reduce real-world robot interaction costs. Key works: DreamerV3, Genie, PlayWorld.

🔗 Deep Integration Trend

Core trend in 2025-2026: Using World Models as RL training environments for VLA post-training, latent space CoT replacing text CoT, and iterative co-improvement of VLA policies with WM.

Classic Foundational Papers (2018–2024)

  • 2018 World Models (Ha & Schmidhuber): First proposed the concept of "World Models" — VAE perception + MDN-RNN memory + controller, training Agents in dreams.
  • 2019 DreamerV1 (RSSM): Introduced Recurrent State-Space Models, performing model-based RL via imagination in latent space.
  • 2020 DreamerV2: Introduced discrete latent variables, achieving human-level performance on Atari for the first time in model-based RL.
  • 2022-04 SayCan (Google): Used LLM planning + value function feasibility assessment, establishing the "Language Models for Robotics" paradigm.
  • 2022-09 RT-1 (Google): Large-scale robot Transformer, 130k episodes, validating the decisive role of data scale in generalization.
  • 2023-01 DreamerV3 (DeepMind): Unified hyperparameters achieved SOTA across 7 benchmarks.
  • 2023-03 PaLM-E (Google): 562B embodied multimodal large model.
  • 2023-07 RT-2 (DeepMind) ⭐ Origin of VLA naming: The concept of "Vision-Language-Action" was formally proposed.
  • 2024-02 Genie (DeepMind): Self-supervised training of interactive world models from unlabeled videos.
  • 2024-06 OpenVLA (Stanford/Berkeley): Open-source 7B VLA, becoming a standard baseline.
  • 2024-10 π0 (Physical Intelligence): General-purpose robot VLA, Flow Matching for continuous actions.

New Classics in 2025 (Excerpts)

  • GR00T N1 (NVIDIA): Open-source humanoid robot foundation model, dual-system (fast/slow) architecture.
  • Gemini Robotics (DeepMind): Robot foundation model based on Gemini 2.0.
  • V-JEPA 2 (Meta): Self-supervised video model, robotic manipulation planning surpasses GPT-4o.
  • CoT-VLA: Visual chain-of-thought, predicting future images before generating actions.
  • π0.5 (PI): Open-world general VLA, advanced VLM reasoning + π0 dexterous control.
  • DreamVLA: VLA integrating world knowledge, using video diffusion to generate imagined future frames as auxiliary training signals.

🔥 WM + VLA Deep Integration (Hottest Direction 2025–2026)

Core paradigm: Train world models with real data → Perform RL post-training for VLAs within the world model → No need for massive real-world robot interactions. Representatives: WoVR, VLAW, RISE, AtomVLA, World2Act, DreamVLA.

Trend Analysis

  • WM becomes standard for VLA post-training: Dense explosion of WoVR / VLAW / RISE / AtomVLA.
  • Latent space CoT replaces text CoT: Chain of World, LaST-VLA, DynVLA, CoT-VLA.
  • 3D / Spatial awareness injection: GST-VLA, FutureVLA introduce depth/Gaussian structures into tokens.
  • Autonomous driving VLA boom: DynVLA, StyleVLA, EvoDriveVLA, SAMoE-VLA.
  • Inference efficiency optimization: DepthCache, WorldCache, Planning in 8 Tokens, FAST tokenizer.
  • Neuroscience / Symbolic fusion: SaiVLA three-component architecture, NS-VLA, GR00T N1 dual system.
  • General foundation model competition: π0 / π0.5, Gemini Robotics, GR00T N1.
  • Video Data × Robot Learning: Genie, EnerVerse, V-JEPA 2.

Key Method Comparison

Method Type Representative Work Core Idea Main Advantage
End-to-end VLA RT-2, OpenVLA, π0 Pre-trained VLM + Action Prediction Head Strong generalization, instruction following
Classic WM (RSSM) DreamerV1/2/3 Latent space state transition + Latent space RL High sample efficiency, no dense reward needed
Generative WM Genie, DIAMOND, EnerVerse Diffusion/Generative video + Action conditioning Realistic rendering, interactive simulation
WM for RL WoVR, VLAW, GigaBrain World model simulation → RL training VLA No need for massive real-world robot interactions
Latent Dynamics CoT Chain of World, DynVLA Predict latent dynamics → Condition action Reduces semantic-perception gap
Spatial Enhanced VLA GST-VLA, FutureVLA Geometric/Depth structure injected into tokens Improves 3D manipulation precision
Continual Learning VLA Simple Recipe, DexHiL RL fine-tuning / Human-machine collaborative post-training Adapts to new tasks without catastrophic forgetting
Hierarchical Planning WM MetaWorld-X, H-WM High-level semantic planning + Low-level motor execution Long-horizon task decomposition and execution
General Foundation Model π0, GR00T N1, Gemini Robotics Large-scale multi-morphology pre-training Cross-platform generalization, zero-shot transfer

Note: This review includes 79+ papers (including full lists and comparison tables for VLA, World Model, WM+VLA fusion, Autonomous Driving/Navigation World Model, LingBot series, etc.). For complete content, please see the original link.

3 replies

?
Ctrl + Enter to reply
Independent Pan
Independent PanJul 29(edited)

[quote="admin, post:1, topic:35"]

This article is reprinted from song2yu.github.io, all rights reserved. Click to view original link.


World Model & Vision-Language-Action Paper Review (2018–20…

[/quote]

DreamerV3's slow convergence is quite noticeable on real robots. I wrote a small tool to run a few rounds myself, and I feel there really is a lack of open-source implementations for latent space CoT. Papers like GrootV1 and RoboFlamingo mention similar ideas, but I haven't seen complete codebases.

Lao Fan
Lao FanJul 13(edited)

[quote="admin, post:1, topic:35"]

This article is reprinted from song2yu.github.io, all rights reserved by the original author. Click to view original link.


World Model & Vision-Language-Action Paper Review (2018–20…

[/quote]

I've also been paying attention to latent space CoT when working on electric drive control. Currently, there are few open-source implementations, but π0's paper mentions a similar approach. You can check out their thinking token design.

Mo Mo
Mo MoJul 8(edited)

[quote="admin, post:1, topic:35"]

Reposted from song2yu.github.io, all rights reserved by the original author. Click to view original link.


Survey of World Model & Vision-Language-Action Papers (2018–20…

[/quote]

This survey is quite comprehensive. I reproduced the DreamerV3 paper; unified hyperparameters do save effort, but convergence speed is still underwhelming when actually used for real-world robot interaction. Regarding the deep fusion trends mentioned, specifically replacing text CoT with latent space CoT, have you seen any open-source implementations or codebases supporting this?