How can a 30B-parameter model activate just 3B parameters per token and still draw on the full model’s capacity? Learn how dense and MoE models use parameters differently, and what that means for throughput, memory and serving complexity. Check out our new technical explainer: ▻ ↧ Dense vs. MoE Models: Active Parameters, Throughput, and When to Ch... How can a 30B-parameter model activate only 3B parameters per token, and still use the capacity of the larger model? Nemotron 3.5 Lightning illustrates the answer: It uses a Mixture-of-Experts (MoE)…