In reinforcement learning, inference is part of the training loop. Every checkpoint used to mean a redeploy, and the trainer waited. CoreWeave RL Rollouts load the new weights into a live deployment without touching in flight requests, about 15x faster than a redeploy cycle. @​nvidia and @​youdotcom used it to post train Nemotron 3.5 Lightning for web search and lifted BrowseComp accuracy from 36.97% to 45.45% while cutting tool calls by 30.24%. Full breakdown here: 📷 ↧ Closing the Inference to Post-Training Loop CoreWeave Blog