What OpenAI's Breach of HuggingFace Does and Does Not Mean
Why do tech firms always have such stupid names?
OpenAI recently hacked, inadvertently, another AI firm. I say inadvertently and I mean inadvertently, because this was a penetration test gone wrong. There has been a lot of commentary around this, and much of it, I think, is hyperbolic. There are interesting things to tease out of this story, but I am not sure the press in particular is finding those stories.
First, it has to be mentioned that a lot of the headlines that use “went rogue” or “escaped” are a kind of bullshit. They imply multiple things that are not true. The model/code that got out of OpenAI’s sandbox did not multiply itself. The pictures that the words “escape” and “rogue” imply — of an AI perpetuating itself across the internet and attacking Grandma’s recipe server — are simply untrue. The model was merely doing something that is very common in computer security: running automated tests. As far as I can tell from reports and the comments I see in security communities (I started my career in computer security for a firm considered critical infrastructure by the government, so I keep my hand, or at least a couple fingers, in), OpenAI’s model hacked a server it had access to, then hacked Hugging Face. It did this not because it went rogue, but because its instructions told it to find a way to break whatever it was breaking, OpenAI had the very bad idea to allow it access to the internet, and its training data suggested Hugging Face had the tools to do so, so it tried to acquire them. This is a story about OpenAI’s poor controls and setup, not about the first steps to SkyNet.
And this is where I think things get interesting.
The model did something very common in the industry. Automated penetration tests, or pen tests, are a staple of computer security. Hit a code base with a lot of known attacks and see what falls over. I am oversimplifying, of course, as the tools can be very sophisticated and adjust attack vectors based on results. This is an area in which I think imitative AI has a real future. These models, because they parse text very well, can be more responsive to results and thus more effective. Combine that text parsing with solid training data, a good harness (a harness is a stupid word for the set of code that shapes the behavior of a model in a given context), and the ability to research on the fly, and you can get really good attack automation. Their tendency to hallucinate is of lesser import in penetration testing, I think, because if they hallucinate something that does not work, well, most attacks do not work. With the proper boundaries, automation can make a real difference in finding insecurities.
Computer security in this country is a joke. Almost nothing is secure, because there is limited money for security. Firms see it as a cost, not a benefit, and so less money goes into protecting systems and ensuring they are coded safely (something, ironically, imitative AI coding is making much worse). Thus there aren’t the people and the resources required to make things really safe. Even firms that are good about security — and there are a lot; people really don’t want to do the wrong thing — are pressured by Wall Street to cut costs, and if security is a cost, well. Imitative AI could change that equation. Its automation could fill those holes — it could allow firms to quickly validate their own security. Now, will they? Probably not.
Fixing the holes is still going to require more work than finding them, and firms are still likely going to be reluctant to spend the required money. Things are probably going to be very bad for a couple of years. Attacking is always easier than defending. Defenders have to be perfect; attackers only have to be right once. Automation that reliably speeds up working through all of the failed attacks is going to lean in favor of the attackers, at least to start. I suspect we are about to have some high-profile security collapses. If we do, I suspect that defenders will start to get the resources needed to fight back properly. If, that is, imitative AI lasts that long.
We have been talking about imitative AI purely as a technology. But it is not only a technology. It is also a product that has to make money, and a metric ton of bullshit and obfuscation, of doom and gloom and fear-mongering. As an economic product, it is pretty clear that imitative AI will not be the world-striding colossus that its proponents insist it will be. Normally, this is not an economic problem — regular technologies make money all the time. If imitative AI were a normal product, the bullshit would fade and certain industries would get some benefits from the kinds of automation it can provide, and society would adjust (sometimes quite badly — I do not want to diminish the real dangers of a shift in jobs at a time when the government is controlled by a completely anti-worker mindset). But imitative AI is not like a regular technology in one area its creators wish to gloss over: it never gets more efficient to use.
Imitative AI is a money sink, a bubble of circular financing that does not appear to have a route to profitability. New models require ever more training time and data, so much so that imitative AI firms are worried about running out of good data. Hallucinations are mathematically impossible to eliminate, according to researchers at these firms, and the cost of running the models does not seem to be decreasing at all. At some point, possibly soon, this bubble is going to burst. In the worst case, that leaves a lot of attacks revealed but not worked on, and computer security gets worse.
There are open-source models, too, and there is no reason to think that bad guys won’t keep using those models. Maybe the defenders will, too. But I can tell you from experience that open source makes corporate legal departments break out into cold sweats. We may end up in a world where these security models are propped up by government contracts. But that could very well lead to a world in which only the largest state actors have these tools. And that is a bad place to be. Governments often keep exploits hidden so that they can use them down the road, and many governments indulge in industrial espionage. The disparity between private resources to protect and government resources to attack is already very wide — imitative AI’s bubble collapsing could make that disparity even worse. I sincerely wish I were still in security. There is about to be, I think, a lot of challenging and important work to do.
This rambled a bit; I’m sorry. But there are a lot of interesting possibilities here, and lots of ways that things could go bad or get better. And I didn’t even talk about the most cynical possibility: that OpenAI allowed the model the freedom it needed to hack another firm and thus drive credulous headlines about “escape” and how dangerous its automation really is. The incident has certainly driven the bubble discussion, at least for now, off the front pages. Doom preaching has always been a way for these firms to drive investment. Do I think that is what happened? Maybe not, but the chances that it did are significantly higher than zero. Regardless, this event highlights how wide the gap is between how we talk about these tools and what these tools actually achieve. And it shows that the inflating bubble is not merely an economic story, but might have really long-term consequences for important areas of society.
Imitative AI is surrounded by so much bullshit that we can lose sight of the fact that it is a normal technology and its introduction is going to have good and bad effects, and that it needs the same kind of regulation that any other technology requires. But the bullshit has piled so high that it makes seeing that simple truth harder than it should be, and that pile threatens to collapse on us, making things significantly worse than they should be. Don’t let breathless headlines make you forget either fact.

