What if making video-based world models smaller isn't just about better compression, but about making their representations easier to predict? Recent approaches to video generation explore building scenes from coarse to fine. The first few tokens capture the big picture, like a person playing a guitar, while later tokens progressively add visual details. Our Interactive Research team just published SemanTok, which takes this idea further. By making those early tokens more semantically meaningful, we give the model a clearer understanding of what's happening in a scene, making the representation easier to predict and video generation more efficient. The result: a model using SemanTok matches or beats the performance of a model more than three times its size. Read the full paper: ▻ ↧ SemanTok: Predictable Semantic Tokens for Efficient Autoregressive ... SemanTok supervises every nested token prefix directly with frozen DINO features, so the first tokens already narrow down the clip's semantics and AR rollouts compound fewer errors.