Trains a world foundation model based on next-state prediction, learning unified latent world representations from video and language for downstream reasoning.