Today we are excited to announce a fresh update on the Lambda Model Cards! Understanding how serving models requires more than knowing the raw tokens/s, and serving real users means caring about more than just concurrency at 1. As of today, our model cards now showcase benchmarks off of @​NVIDIAAI's AIPerf benchmarking package on minimal Lambda hardware. With this, you can now properly get an understanding of how your model deployments will behave at certain user thresholds, what its interactivity will look like, and know when you need to scale up to more replicas. 📷 ↧ Inference models Lambda's catalog of model cards for the LLMs that matter. Search by model name to get architecture breakdowns, hardware requirements, deployment guides, and throughput benchmarks on NVIDIA GPUs.