New research: Training a Misaligned Reward Seeker What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable. In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring. Read more: 📷 ↧
published
New research: Training a Misaligned Reward Seeker What produces severe misalignment? We’ve…

Latest documented BWB result
See Billy's timestamped results and follow-through.
Review the latest gains posts, original timestamps, and proof images published by Banking With Billy.
Historical results are not a promise of future performance. Trading involves substantial risk.This public post is a timestamped information archive, not personalized financial advice. Alerts can change as markets move. Join Billy's private group for the complete daily stream and follow-through.
