A video world model can make a convincing clip and still get the physics wrong. Our researchers just released Physis-Lang, an open self-evolving framework that adds physics reasoning to video captions. The captions explain why and how a scene unfolds. We use them to fine-tune world models and add that reasoning to prompts when generating video. Adding physics reasoning to the prompt alone improved NVIDIA Cosmos 3’s PhyGenBench score by 5.62 points without retraining. Paper and project: ▻ ↧ Physis-Lang Self-Evolving Language for Video World Models Physis-Lang makes physical language a shared, self-evolving representation for data curation, model training, and video generation. Research from NVIDIA.