A realistic video is not enough for a robot to act reliably. That was the central question at CVPR 2026’s WMAS workshop. RoboWM-Bench, the Best Paper, found that a video world model can produce convincing footage while still failing when its predicted behavior is turned into robot actions. The failures came from problems such as spatial reasoning and unstable contact. GEM-4D improved real-world manipulation success from 61% to 81% by adding geometric consistency across generated frames. SAW-Bench found a 37.66-point gap between humans and the best multimodal model on situated-awareness tasks. 📷 ↧ From world models to active agents: the next step for physical AI Inside CVPR 2026's WMAS workshop: why world models need active sensing and closed-loop planning to become real physical AI agents.