MLPerf Inference v6.1 is out, and we have two firsts: the first agentic inference workload on datacenter hardware, and the first MLPerf deployment of a model over a trillion parameters. On 4x NVIDIA Blackwell Ultra GPUs, we posted the leading Offline throughput on GPT-OSS 120B among all Blackwell Ultra GPU submissions, plus an 8.85% throughput gain over v6.0 on identical hardware. That's six months of pure software optimization. On NVIDIA HGX B200, we swapped in Kimi K2.6 for the open-division agentic benchmark, running a 1T+ parameter model where the reference workload expects 27B. Same harness, no memory ceiling. Full results and methodology in the blog. Link below. 📷 ↧