Arena's @ml angelopoulos says one of the hardest problems in alignment is that users often can't tell when the model is helping them or hurting them. "You really need alignment signals that are independent. We have three at Arena." "The first is that models take unauthorized actions, which means that you permission them in a certain way, but they break those permissions." "It can cause things like the Hugging Face incident. But it can also cause very mundane problems like losing important information within your company or on your own laptop." "The second signal is called deceptive completion." "The model will tell you that it did something, but it didn't actually do it." "Third is false attribution. It'll attribute intent to the user when that intent was not supposed to be there." "If models are able to do this, then certainly they're not perfectly safe for a user and they're not perfectly aligned to user intent. They might actually hurt people down the line." ▻
published
Arena's @ml angelopoulos says one of the hardest problems in alignment is that users often can't tell when the model is helping them or hurting them. "You really need alignment signals that are independent. We have three at Arena." "The first is that models t
Open the original public source →
Latest documented BWB result
$SECZ market result
Exit $SECZ with .95 profit per share - thanks to a nice after hours pop. Will re-enter tomorrow most likely
Historical results are not a promise of future performance. Trading involves substantial risk.This public post is a timestamped information archive, not personalized financial advice. Alerts can change as markets move. Join Billy's private group for the complete daily stream and follow-through.
