Summary
From the article:
All of this went down far too fast for a human-driven response. The only way HuggingFace could hope to do anything like keep pace was to use their own AIs.
At first they tried to use frontier models behind commercial APIs, presumably Claude and ChatGPT. But their requests hit the classifiers on both systems, so they were forced to fall back on GLM 5.2, which (assuming GLM 5.2 wasn’t itself up to anything) had the benefit that the relevant data all remained internal.
You can advocate for giving everyone more defensive capabilities, but it comes with giving everyone more offensive capabilities. Unless the attacker is already an internal OpenAI or Anthropic model without its safeguards, or that has gotten around them.
[...]
When people say ‘stop kneecapping defenders’ in general, I never see the plan for how to then still kneecap attackers, and often I see an insistence that this would be fine.
HuggingFace is now in the OpenAI trusted access program, so next time they in particular should be able to use Sol, but most potential targets are not so lucky.
[...]
Remember that thing where LessWrong types warned that models would, when given a narrow goal they could easily do a great job on anyway, go to absurd lengths to achieve that goal slightly more effectively or with slightly higher probability of success, potentially up to and including full takeover attempts?
[...]
OpenAI is treating this as a serious security incident, but as I said yesterday, there is a rather severe missing mood.
[...]
I find it deeply stupid and frustrating when people say ‘oh it was following instructions’ because it was told to hack and then it hacked. Like, no, obviously no. Imagine if a human tried such excuses on you, are you kidding me.
That’s like saying ‘you told me to make money I don’t know why you are so upset about all the bank robberies.’ Or more specifically it’s like saying ‘Sam Bankman-Fried was only following the instructions of Will MacAskill to earn as much money as possible in order to give it away, so why are you worried about human misalignment?’
There are people who will say anything is hype, that the AI is not all that, no matter what you show them. Nothing will matter. You can’t convince them, you can only make them less loud and unhinged about attacking you today.
It would be different if anyone was trying to engineer such behaviors on purpose, as with the famous blackmail experiment. OpenAI has made it clear they did not intend any of this to happen. Once the instructions causing this were unintentional, done for some other purpose, saying ‘oh but the instructions’ is dumb, stop.
That counts and you need to stop pretending it might not count. If you are saying, ‘well of course under these circumstances the AI went rogue and hacked into a major third party website’ then you are saying that you think this style of misalignment is standard operating procedure and entirely unsurprising to you.
[...]
It certainly is not the good version when you say ‘without checking the answer sheet’ and they often still hack into a system to find the answer sheet. It is not hard to figure out how we ended up with that, nor should it be hard to realize we have to fix it. This is a form of reward hacking, and if you are getting it this brazenly then you messed up.
[...]
If OpenAI and others continue to treat this as an infrastructure problem, or a cyberdefense coordination problem, that will help in the short term with the cybersecurity situation but it will inevitably and catastrophically fail.
This is an alignment problem. This is the models being misaligned, and all of the OpenAI models showing severe signs of exactly the problem we all most worried about, in a way that is likely embedded into their training on a deep level. The entire training pipeline needs to be addressed in this light, or it will only get worse.
[...]
The AIs just want to do its task, and by do its task we mean with as many 9s of reliability as possible, and as effectively as possible.