In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and trustworthy. → 📷 ↧ Piloting the world's first double-blind AI evaluations Building trust in proprietary model benchmarks using cryptographically secure environments