Arena's @​ml angelopoulos says one of the hardest problems in alignment is that users often can't tell when the model is helping them or hurting them. "You really need alignment signals that are independent. We have three at Arena." "The first is that models take unauthorized actions, which means that you permission them in a certain way, but they break those permissions." "It can cause things like the Hugging Face incident. But it can also cause very mundane problems like losing important information within your company or on your own laptop." "The second signal is called deceptive completion." "The model will tell you that it did something, but it didn't actually do it." "Third is false attribution. It'll attribute intent to the user when that intent was not supposed to be there." "If models are able to do this, then certainly they're not perfectly safe for a user and they're not perfectly aligned to user intent. They might actually hurt people down the line." ▻