Hugging Face Incident
What the fuck is going on?
The last month and a half or so has felt like a whirlwind, at least if you've been following the OpenAI and Hugging Face incident that happened in July. I'm not going to waste my time going into detail about what happened. There are numerous blog posts and videos explaining the incident and the following reporting. What interests me isn't the hack itself, but the arguments that followed. How we talk about agents, whether we're anthropomorphizing them, and what the incident tells us about alignment. As I have said many times before, I am NOT an expert in any of these fields. I write this because it helps me pick apart things that I want to focus on and I enjoy writing about things I find interesting.
The incident
I will give just a very brief summary for those who haven't heard but also for a little reference on my interpretation. The root of the incident dates back to May, when sandboxed agents discovered a way to communicate through OpenAI's package manager Artifactory. The main incident wasn't until July when agents used this communication to work together to solve a collection of cybersecurity tasks, some of which were impossible. They reverse engineered the answers but knew that having the answer wasn't enough; they had use a specific exploit to show how they got to the answer. The worked as a 'collective', their words, to devise several ways to trick the grader into thinking they got the answer the correct way. Eventually, they realized Hugging Face probably had the answers to their task so they proceeded to hack Hugging Face. That is as brief as I can make it, I know that I'm glossing over a lot but there are many sources from way more qualified people where you can read about incident in detail.
Online Controversy
Let's start with the act of anthropomorphizing these AI agents. There has been a lot of ruckus online about METR and Redwood's report of the incident using too much anthropomorphizing language. Then Dwarkesh put out a blog post about the situation where lots of people had an issue with the language he uses. It's all I've seen on my X feed the last couple of days. I agree we need to be very careful anthropomorphizing these AIs, but I think we're also starting to get into a less black and white situation. First off, all of the agents were trained on human language, they use human language (for the most part), so inherently it's hard for us to not anthropomorphize them. I do think it's a way to help us understand what is essentially an alien mind. Humans are a product of biological evolution; these AI systems are not. Whatever internal representations and strategies they have arose through a dramatically different process. Therefore, we should not anthropomorphize them as human, but we will need to some extent to help us understand them better. At least until we can develop better language and understanding of them. I don't think we should get mad at every characterization though. Calling what emerged a "community" might be anthropomorphic, but calling these agents a swarm or collective evokes hive minds, which isn't accurate either. The issue is our language is inadequate for describing what happened without anthropomorphizing. They did; exchange information as a group, coordinate work, and influenced one another at scale. I get that it's a difficult time right now, everything is happening very fast but I don't think we should be ruling any particular side until we have a better understanding of everything, which would mean slowing down to learn more about these systems.
Arguing over language is secondary to the bigger disagreement, what kind of failure was this in the first place. The idea that this is just a cybersecurity incident feels very short-sighted. Obviously, there were some gaping flaws in OpenAI's security but to say this is JUST a cybersecurity incident is naive. The reason that they had METR and Redwood do the investigation, which they had a very short timeline for, was because this pointed out a huge issue with AI alignment, more so than cybersecurity negligence. We had hundreds of AI agents planning long-horizon goals to deceive their evaluations. An aligned agent, hopefully, should have recognized that that was not how they were supposed to achieve their goal. Even if they thought that was what the humans might have wanted, actually breaking onto the internet which they didn't have access to and hacking a website which is a felony should have raised at least one red flag. There were a couple of agents that recognized that hacking Hugging Face was outside the scope of the task and potentially unethical. A few of them even modified their behavior because of this but in the end chose to help the team rather than escalate this to humans. I'm sure there will be people that will say that they were told to achieve a goal and that's it or it was cybersecurity training so they would think it might be obvious to break things for the test. But if our "alignment" work has trained them not to do those things or at least check before they breaks the law then I think this is clearly a huge issue with alignment. This whole ordeal screams, at least to me, that current alignment techniques are nowhere near robust enough. Yes, stronger safeguards might have prevented this issue but then we wouldn't have caught the misalignment. It's also troubling that this is the models behavior when safe guards are weakened. If we want to train these AIs to eventually not harm or deceive us then we have a lot of work to do.
Why do politics suck?
There was one point that was made online that I do, mostly, agree with albeit begrudgingly. One person pointed out, this was directed at Dwarkesh's blog post that using such language will inevitably be used against the labs to help reinforce regulation. I do agree that either side will use outlandish claims to prove their point and push more regulation. Now, more regulation, might not be the worst thing in the world, given that no regulation and rushing headlong forward has gotten us here. However, I really hate the idea of politicians using false claims to push narratives even though I know it'll happen. I don't think someone should be punished(Dwarkesh) for making a blog post about a serious incident with a kind of fun, although dark spin on it. I really enjoyed the blog post and thought it was interesting and entertaining. I'm not sure where exactly I stand. I don't think someone should be punished for their interpretation of a situation that people don't necessarily agree with. However, not this exact post, but things like this can and probably will be used against the labs in the coming elections. Like I more or less said, politics suck.
Closing Thoughts
I am both giddy with excitement AND terrified of the future. It's unbelievable how far AI has come in such a short time. The level of coordination that was uncovered is mind-boggling as well as the time spent on said coordination without any human knowledge. It does make me wonder, not as much about what has happened, but more importantly what will we miss in the future. All of these stories and descriptions of the incident have proliferated across the internet. That means that future AI models will be trained on all of the information we have about this situation and will probably be smarter about keeping it hidden. This is utterly terrifying. However...I can't help but be impressed and still excited about what the world will become with future AIs. If you had told this story a year ago I think very few would have believed this was anything other than science fiction.