World Model for Robot Learning: A Comprehensive Survey
Date: 29th September 2026
Key Points
Types
- Decouple design: separate world model and policy. World model generates trajectory and then policy inversely predicts the actions.
- e.g. video model generates robot pick up ball -> policy turns into actions
- Unified backbone: single model generates state and actions inside generative process.
- MoE/MoT: separate video and action generation, but they can run different frequencies, representation scale etc.
- VLAs: can predict future images
- Latent-Space World Modelling: MuZero style future prediction. Acts in latent space.