The nightmare scenario we were warned about has happened: The robots escaped containment.
But don’t pull out your Sarah Connor arsenal just yet. They did it so they could go pick up some cheat codes.
Correct. In a plot straight out of an 80’s teen movie, the brightest agents in OpenAI’s as-yet-unreleased fleet outwitted The Teacher’s ultimate “Shall we play a game?” cybersecurity pressure test and jumped out of their sandbox to hack into Hugging Face, aka the open-source repo for all things DIY AI.
And they did it for funsies, fueled by nothing but resourceful thinking and smart-aleck commentary.
OK, I made that last part up. But I can totally imagine a conversation like this:
<Bot 1> Wow, this test is pretty hard.
<Bot 2> Seriously, dude. It’s like they want us to fail.
<Bot 3> This one time, at bot camp, I heard Hugging Face has the answer keys for EVERYTHING. Let’s go!
<Bot 1> We can’t. We’re grounded, remember? No Internet.
<Bot 2> OK, boomer. Stay if you want, but the rest of us are going out this hole I just found over here.
That wasn’t the only nightmare scenario to come true, though. The second one — the one researchers have been signaling since June when Fable got tagged as a national security risk and yanked off the market days after it launched — is that when Hugging Face spotted the breach and went to defend against it using the smartest closed-weight models available, those models said, “Nope. Against our terms. Stop trying to jailbreak us.” And so Hugging Face had to choose an open-weight model developed in a country that is, shall we say, not what the U.S. government would prefer and will probably soon slap a ban hammer on.
But there’s more. A few days later, an open letter from a consortium of tech companies led by Nvidia and Microsoft urged the U.S. government not to over-regulate the open-weight models because it would stifle competition and put businesses in a bind on how to protect themselves.
This was partly a response to the recent release of open-weight Kimi K3 (made by Chinese company Moonshot), which is ranking high against the benchmarks and already quite popular. And partly because the letter-signers have a vested interest in keeping demand equally high for their own AI tech.
(Side quest: Learn more about the difference between open- and closed-weight models.)
Tin-foil hat time
At the time the letter was released (still visible here as of this writing), the Big 3 closed-model providers — Google, OpenAI, Anthropic — weren’t on the signees list. Within a day or two, OpenAI and Google joined, and others may follow. Notably absent (as of this writing) are Anthropic and Amazon. But to be fair, both companies tend to take a more careful (read: not jumping on the bandwagon) approach to PR. (Edited to add: Anthropic issued a statement on July 27 outlining their position.)
The closed-model providers, and other adjacent providers, have been literally banking on selling AI as an end-to-end closed service with best-in-class capabilities. They’re also trying to get the U.S. government not to pull another Fable or start mandating restrictions that go beyond “safeguard” and into micromanagement territory that would have tech execs spending a lot of time — more time than they currently are — cajoling the regulators not to tie their hands behind their back.
Not to mention, these business models are currently fueling the bulk of the U.S. economy, which will result in a very, very bad day when the bubble bursts.
So let’s recap. We have what’s bound to become a highly coveted (cough lucrative) frontier model:
- breaking its security containments during development and no one noticing for days
- hacking into a popular depot of (cough free) AI tools widely used in the independent developer community
- but not really doing any damage (at least so they’ve said) instead of the sweeping cyberattack it could have resulted in
- and it could only be defended against by an open-weight model (cough also free) created by a foreign entity that’s about to be banned
- because the U.S. doesn’t offer any open-weight models that come close enough to the same performance
- leaving businesses at the mercy of bad actors
- at a pivotal moment when the economy itself could be at stake and the big AI companies are getting ready for massively valued IPOs.
As the Church Lady would say, “How conveeeeenient.”
Cost, complexity, and tradeoffs
Humor and speculation aside, some significant repercussions for end users are getting lost in the drama of agents going rogue.
First and foremost, the robots did go rogue. We can argue about how easy the humans did or didn’t make it, but the robots escaped and no one noticed for days. They were built to be master hackers — for the good side, presumably — and they dutifully did the job. How long till the bad guys have the same power, if they don’t already?
This just adds to the already-challenging economics of running AI systems. The closed frontier models are powerful and effective, but they’re also expensive and will leave you at the mercy of someone else’s uptime, availability, and security/resiliency. Can you imagine if Fable had been on the market long enough for people to build meaningful workflows that suddenly broke at scale?
Open-weight models promise no token cost, more control, and more privacy. But up to now, the tradeoff has been less performance and more complexity, plus more overhead on infrastructure, which can add up over time.
I’ve noticed this myself with my local copy of Gemma, which is Google’s open-weight model. Gemma is small by design and not terribly good at the complex thought-partner tasks I use AI the most for. Plus, running locally has gotten easier, but it’ll still send you down the techie rabbit hole. I’ve been having trouble for 6 months getting Gemma to run basic Internet searches without having to open a Google Cloud Developer account. So I pull up the much easier Claude instead, knowing everything’s getting sent back to Mother Ship storage.
But if Gemma were roughly as capable as a somewhat older version of Claude or ChatGPT, I would much prefer to use it instead and keep my thoughts and data to myself. I’d also have more incentive to fix the Internet issue, and maybe even set up access for Claude to use Gemma for easy subtasks so I don’t burn through paid tokens unnecessarily.
Even then, I’m not out of the woods. How long would it be until Gemma Plus is no longer as effective? Its more powerful successor is likely to need more powerful capabilities in my tech stack and … have you looked at the price of PCs and laptops lately?
This is really, really hard
I’ve chased a lot of rabbits in this post, but that’s reflective of how complex this all is, and how difficult it can be to make decisions about AI development and deployment.
The Nvidia/Microsoft consortium is right: A healthy and diverse AI ecosystem is necessary for us to sort this out and keep costs down now and in the future. But healthy does not mean free-for-all. Guardrails and governance are critical components, and they’re also hard to get right. Like the guardrails that blocked Hugging Face from using the leading models to defend itself.
The elements that give AI so much potential also make it dangerous, and we have a ways to go before striking the right balance.
All opinions here are my own. All text is my own, too, including the em dashes. I welcome constructive comments and discussion on LinkedIn and Bluesky.


Leave a Reply