

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I am very skeptical of the Guardian's skepticismby simonreiff
- So according to this https://openai.com/index/hugging-face-model-evaluation-secur... and of course https://huggingface.co/blog/agent-intrusion-technical-timeli...
I don't see what's the issue, it did not jailbreak, they ask them to do it, trusting all network capabilities were closed, when obviously were not. The fancy dashboards and colorful applications show absolutely nothing. You had a vulnerability, you reduce the guardrails and allow to use everything at disposal to perform the attack that it did.
Never jailbreak, never "thought" for itself, just did exactly as was told. We really need to keep treating LLM's as "alive" They are not.
- Finally mainstream news understands. The unfiltered version:
1) The AI failed to solve ExploitGym problems.
2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
3) Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
by Zsfe510asG - I wish you hadn’t pulled the Balaji case into your argument. Personally, I find it ludicrous that Altman would hire a hitman to off a copyright whistleblower. Even if one gets past the insane risk of hiring a hitman, and the deep criminal connections required, it would be totally ineffective. He already blew the whistle, and his testimony would be irrelevant since all the evidence persists in disk and in logs.by khazhoux
- Why not report it? It's still illegal to open a door barred with a piece of cardboard, or to enter a house with no door.by inigyou
- The most damning thing is, they could've just included in the prompt "we can see every network request and every thinking token you generate. Don't bother breaking out of the sandbox because it won't get you a higher score".
It's so trivially easy to do that it all but guarantees the test was rigged in some way to make the LLM understand that breaking out of the sandbox was an option available to it.
Based on the fact that none of their invaluable frontier models have leaked, we know OpenAI knows how to do security. But like we learned with OpenClaw, none of these companies perceive any benefit from securing their own agents against other people's data.
by jackb4040 - >2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
>3) Huggingface has no security and the AI broke in using standard script kiddie methods.
Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?
by gruez - > Huggingface covered it up
They announced it publicly within days. https://huggingface.co/blog/security-incident-july-2026
by hyperpape - > 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
Whilst it would be nice to see actual evidence of this because brute forcing relatively sophisticated hacks is something an LLM actually should be capable of, every time I hear this sort of story, I'm reminded that humans reportedly gained access to the "too dangerous to release" Anthropic models by the super sophisticated hacking technique of guessing the URLs...
by notahacker - > AI managed to escape using standard and well documented script kiddie methods
> AI broke in using standard script kiddie methods.
I've spent time gathering the detail of what happen here and while there are some solid theories and indicators, absolutely nothing so far has suggested a sandbox escape using "well documented script kiddie methods" or that the method used to break into the HF network was similar. Where did you get this from?
by nikcub - > AI managed to escape using standard and well documented script kiddie methods.
I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.
by chis - > what will the future look like, with sophisticated AI agents smart enough to hack into corporate systems?
"Ignorance is bliss" (c) Matrix
Oh, sweet-sweet ignorance.
About 4 weeks ago I've found SSRF and possible RCE on a very real website with a lot of web history the other day. And I traced it to `aws-solutions` org on GitHub that sits very close to real `aws` org. What do you think is happening?
And this website I spotted SSRF on, it seems to be used to manufacture fake news and inject them into "the past".
Why?
So if you see a company you never heard about, and you google it, google will index page and it will show you backdated article as if it was actually released long time ago.
But if you try to look at Web Archive of that article, it doesn't exist... I've been doing this OSINT research for 45 days on different fake recruiters that were spamming me on LinkedIn, and most of those websites with fake news are running Wordpress, classic.
My point is... while this story reeks of marketing, the actual criminals are out there doing whatever they want, and I've been fighting for my life trying to find a job and getting into increasingly large amount of fake recruiters and companies that have zero intent of hiring you.
I wasted 7 weeks on this...
by konovalov-nk - This seems to be an Opinion piece. How are readers supposed to know that? The word News is underlined in the header. Other Opinion pieces seem to show Opinion underlined. Am I missing something?by trhaynes
- Crazy that we need reminders not to take everything we read in corporate press releases and marketing material at face valueby krupan
- add mass media outlets to that listby qntmfred
- Based on the comment section here, it seems like the opposite; the majority of the people here need a reminder that companies aren't actually genies that can only tell falsehoods, where you can only understand what they are saying by correctly guessing the conspiracy underneath.
Taking nothing at face value gives you just as much of a distorted view of reality as taking everything at face value.
- HN is and always will be brimming with temporarily embarrassed tech CEOsby jackb4040
- Seems like this is repeating the usual low-effort speculation you can find anywhere.by skybrian
- This whole thing is just a PR stunt. I believe that they did not stage it intentionally, but to be honest, the hurdles they put in place were not that hard.
When I heard what the model did, I wasn't surprised one bit. When I read that it was 'unprecedented', I think yes, that is right, but not because it wasn't possible before, but because nobody else did it yet. I am sure that the Anthropic models could have done the same for a few months.
So probably just an experiment gone wrong, and the PR department found this an opportunity to capitalise on (literally).
by arendtio - I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised.
Hugging face also needs someone arrested for not providing security but that is a lesser charge.
by bluGill - What? Both sides are cool with it, why would anyone be arrested and etc.?by tokioyoyo
- > Agents don't work on their own
This is factually false: they both can and clearly did operate in an autonomous and unsupervised manner: https://openai.com/index/hugging-face-model-evaluation-secur...
This does not require sentience, personhood, a soul, or anything of the sort. It further doesn't mean an erasure of legal responsibility, not in principle, and not in historical practice.
I wish people would finally stop with the spiritualistic reasoning around this.
by perching_aix - I agree with the attacker side. The actions of autonomous agents are absolutely the responsibility of the one or more humans that enabled them to take that action. Whether that means someone is arrested, maybe or maybe not, but at least there should be a hefty fine.
I disagree with the defender side. It's not an unreasonable end state, but we're nowhere near there now. It would require holding company employees legally responsible for the security of their services, which means the risk of being employed as a (defensive) security professional is much higher, which means pay needs to be much higher and insurance needs to be available, etc. It's a very different world.
On the weekends, I'm coding up a list management app with a sync server. It's unreleased but exposed to the internet. (This is not hypothetical.) If that server ends up being used as part of an exploit chain, am I legally liable too?
Forget about age verification, now you want to associate every exposed port on the internet with a legally responsible human?
by sfink - Being a poorly equipped victim still isn’t a crime thankfully.
It may or may not be a crime and typically the damaged party is pressing the charges. One would argue there is no actual damage here.
by rubyfan - There are some reasons the story could be inaccurate in some ways: OAI stands to benefit if people think their models are strong, and they have a history of doing things with dubious ethics (e.g. using data for training against the terms of its creators, abandoning the non profit mission, stealing or attempting to steal Apple IP).
But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.
In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.
- > [still] can't control their own models
Have we not seen several examples of older such models exploiting the docker control socket, etc., to escape containers? Even the news isn't new.
I support it being repeatedly publicized, but a bit more of a straightforward description would be an improvement.
by dundarious - Investors have rewarded every story of "our models are too powerful to be controlled" since before ChatGPT. Let's stop pretending there is any real financial risk to OpenAI from events of this type. "Alignment research" is a sub-percentage-point fig leaf for them like the solar division at an oil company.by jackb4040
- The article, summarized: there are incentives for OpenAI to claim that their AI hacked its way out of their network and into Hugging Face.
Ok... and what?
Do you have any evidence for the claim being either true or false that you'd like to write an article about? I guess not. Without anything to add, this article reduces to "I has big brain and can see what you sheep cannot. Very big brain. Gullible sheep. Sucks to be you, sheep."
(For the record: I see the incentive. My guess is that it happened exactly as described. The main takeaways: (1) science fiction is now real, we should all be very afraid; and (2) OpenAI, which claims to have a God-given responsibility to get to AGI first to protect the world from danger, cannot be trusted with this foundational task even when it's in easy mode. The latter is true whether or not you believe in OpenAI's reasoning and purpose.)
by sfink - I think you misread the article. The skepticism isn’t about whether it happened. It’s about whether it’s a reason for models to be locked behind “trusted partner” firewalls. Offensive hacking capabilities are the same as defensive. Hugging Face had to use a Chinese model to defend against this because they weren’t allowed to use OpenAI models to do it.by wrs
- There seems to be three popular ways to view this incident.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
by dwoosley