Summary
From the article:
Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. And also further kudos for actually taking the model offline for a time to build new safeguards. They gave us one hell of a candid report.
[...]
It was not shared on the official account because OpenAI worried about it being seen as self-promotional hype. It is crazy that one needs to worry about that, but also plausibly a real concern. So again, good decision.
Not that any of the behaviors or failures here are unexpected, exactly. Not by the AIs and not by the humans. Yet there is something I would call a missing mood, a failure to realize the gravity of the situation.
There are some who responded ‘what part of this was unexpected, exactly?’ And that is actually fair, but that is also the problem. We have become numb to all this. We expect the models to be misaligned, and for us to respond only insofar as this presents a practical issue with currently proposed deployments.
AI control is a fine defense-in-depth strategy, as is reducing frequency of practical incidents with things like better instruction remembering. I am very happy that OpenAI is making an attempt at AI control here. I want to be clear that, centrally, OpenAI has done a good thing, both by pausing internal deployment to build new safeguards, and by telling us about this in detail.
But if your models are fundamentally misaligned in that they will, when feasible, use early forms of instrumental convergence to complete the assigned task even when this involves circumventing their instructions and restrictions and is obviously not what the user wants or should want - the most classic alignment failure of all, the stuff of The Genie Knows, But Doesn’t Care and The Hidden Complexity of Wishes - and you know this, I do not accept ‘we will monitor them and catch their constant escape and hacking attempts as they get better at doing so’ as a medium or long term solution.
[...]
The solution is not alarmism, but the correct amount of alarm is not zero.
That, and recognizing this as a Total LessWrong Victory, in the sense that everything is going as predicted, and also a Total LessWrong Defeat, in the sense that everything is going as predicted.
[...]
They trained the model to keep working on its own for long periods of time without giving up. The model interpreted this persistence as including, when it hit the limits of its sandbox or other environment, trying to escape the sandbox or exploit the environment. Whoops. The models they just want to do tasks, you know?
[...]
Saying ‘the model was unable to find a vulnerability’ does not make this okay. If the model knowingly is looking for vulnerabilities in order to escape, and this does not even surprise you all that much, something has already gone terribly wrong, and you know what you should expect when the models get smarter.
[...]
If your model is suddenly finding multiple surprising ways out of your sandbox, then you should expect future smarter models to find many additional surprising ways out of your sandbox, even after you plug the current holes.
The seemingly amazingly great news is that all these models are largely escaping in order to notify us that they have completed requested tasks. For now, we believe, they are not being strategic enough to realize they should not be blowing their cover on that. So we get to notice that the AIs are strong enough that, when sufficiently motivated, they can increasingly get out of sandboxes.
[...]
The correct response to ‘the model keeps trying to circumvent the system’ should be the same reaction that you have to ‘a person keeps trying to circumvent the system.’ Which is that you need to lock them out of the system entirely. Not only here, but permanently. They’re fired. You lose. Good day, sir. Misaligned.