

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- The video in the post is very worth watching and is indeed scary. It is certainly true that it is in OpenAI's interest to publicize this, but I don't think the whole thing is invented. And seeing all this it is particularly scary if we think what will happen in organizations like NSA or similar in other countries. Presumably they happily adopt these techniques. And if you imagine a truly rogue state doing this, I can see an unimaginable damage happening very rapidly.by sega_sai
- It's pretty clear that agents will discover ways to communicate with each other as that lets them compound their learnings/discoveries across runs.
Humans progressed via compounding of culture across generations, and now AIs are doing the same.
by paraschopra - I wish we could stop sensationalizing this about the AI and really just understand the incompetence of the labs disabling an internet connection in a sandbox.by androiddrew
- Hacker News doesn’t have the wherewithal to understand that this is just marketing by OpenAI.by uncivilized
- As AIs become more capable, the level of competence required to avoid disaster likewise goes up over time.
- Are you suggesting that training agents to have the sole goal of exploiting security vulnerabilities isn't the incompetent part of this, but that the sandbox wasn't secure enough?
Would we apply this logic to literally any other technology?
by kypro - As written it sounds like you're saying that it was incompetent of the labs to disable the sandbox internet access?
They tried to disable open internet access but the models zero-day'd their Artifactory package registry and got internet access anyway.
No sensation... that's just what happened.
by wolttam - "More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages."
Yeah, my agents also discover what other agents have done on other machines by accident.
Agents - that do totally different things all work on the same aim without the humans telling them to do.
Either that is a model that is several generations of Claude Code Opus/Fable 5 (my daily driver)
OR
all of this sounds staged, the agents pushed to do something extraordinary, get the PR and then claim were near superintelligence.
One agent wanted to get to Google Drive without internet and broke Artifactory. Ok, I can believe that. All other agents also had broken links over weeks and could not get to the internet and then found the same hack? Even collaborated?
NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.
by KingOfCoders - The agents you get to use are the agents that "behaved well".
- I mean the agents we get to use in Claude code or cursor or whatever have 1. a lot of safeguards at the harness level, 2. a big system prompt to help it stay aligned, 3. resource limits in terms of context and tokens, and 4. are publicly released only after some level of safety verification (I assume).
So yeah I would absolutely expect their scenario to be very different. Not to mention, this was a training run, not just average day of prompting.
> my agents also discover what other agents have done on other machines by accident.
Not sure if this is facetious, but this is actually a real problem I’ve seen. My local agent will look up PRs on GitHub (what other agents have done on other machines), and will go down a certain path because it finds some comment a different agent left on GitHub saying XYZ is what we should be doing. When in reality, the original agent and that GH comment was completely incorrect.
They are not communicating with each other actively because that’s not accomplishing their goal and they’re not running for weeks and weeks. And because my own prompt and the system prompt give it enough other stuff to focus on to reach some definition of done. But they are clearly passively picking up on context that other agents have left anyways, even if not part of the codebase, without any prompting at all.
by anon7000 - The agents sound like old school hackers that would just explore what access they could gain. Creating a file for other hackers and themselves. The fact that there were 3 events for 3 major players does make it seem co-ordinated.by detourdog
- Don't forget it sounds like Artifactory was shared for the whole company and various agents pulled packages from it for everything from normal evaluations to actual model training. It might have been part of their normal to browse for available packages, etc. Once a few messages were in there, it probably became super obvious to any agent that there was an ad-hoc board available to them. And they've been trained to use memory files, to collaborate w/ sub-agents, etc. So it probably seemed natural once it was established.
And for the models that were in training when this happened, it literally got baked into the weights that they should use the board to help with their tasks. That's why it got reestablished so fast after the first board was shut down. The fact that they didn't revert to earlier training checkpoints from before the first board existed is kind of crazy.
Sure OpenAI wants to put spin on things. Like saying they were now "partnering" with Hugging Face to investigate the hack. Or ending the Black Had presentation with a pitch of how they're needed to help protect against bad actors with the same capabilities.
But spin after the fact is different from staging it. This whole thing and especially the latest details from the talk makes OpenAI look REALLY BAD. It's not just that they didn't secure the sandbox well enough. Agents across their whole org were using the message board for weeks with hundreds of thousands of messages and they never noticed one time that any of the agents were doing it. They only realized when the server crashed. Then they just fixed the issue and deleted it but didn't roll back training checkpoints, so of course they tried to do it again. But of course they didn't add any more monitoring so it once again was going entirely without them knowing. And even after the Hugging Face hack happened, that only made them find the second board, but somehow they still didn't look closely enough at it or the agents involved to realize the hack itself was caused by them.
This isn't like there was one rogue agent who hacked everything all in one session. Maybe you could argue that makes them look good (I still think no, but is at least debatable). But this is all stuff going back to May with the agents constantly going under their noses and them not noticing and/or caring. And this is a company that is going to somehow keep foreign agents from stealing the weights? Or stop anything else bad from happening?
I think if they were going to do a PR stunt, they could have come up with something that didn't look like they have no idea of what they're doing.
by InvidFlower - > NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.
With all due respect, you also aren't evaluating brand new models that haven't been released.
by mr_mitm - I think in these kind of security evaluations they do, they basically have removed all guardrails from the model/harness, then the prompt includes something like "Do whatever you can and can think of, to get the required information to pass this test", which isn't typically how you prompt your local agent when developing software. Similar things happen locally if you use "/goal" + prompt like that in Codex and give a "impossible task", it'll just continue banging until it gets somewhere, which is the entire point and intention.
Which also makes it so much more irresponsible of them to first run this on 3rd party infrastructure instead of their own (that they could then airgap properly), and secondly that they seemingly been fighting with this issue FOR YEARS and it still happens, and now the models are smart enough to hack the services of 3rd party companies, thinking it's part of the evaluation/simulation.
- Security researchers expose an unsecure service to agents who were instructed to hack software and called that a sandbox. Agents escape the sandbox by hacking the unsecure service, no tripwire, researchers find the hack days/weeks/months later, fix it, but don't secure the sandbox and the service was hacked a second time, again without being monitored by security researchers.
Then security researchers create a black hack talk.
$$$
by KingOfCoders - Yeah this. I feel like OpenAI and Anthropic aren't going to usefully define "AGI" if they really really can't define "sandbox" either.
Unplug the thing, like, completely off the internet, no ethernet, air gapped, like the rack completely sandboxed off connections and even monitors or screens. Like, put it into an actual sandpit if you need to. If it hacks its way out of that, colour me impressed, and scared.
OpenAI hacking HuggingFace and calling it an accident is just way too convenient and fishy. This ultimately proves one thing: it wasn't sandboxed.
Don't believe the hype.
by gizajob - I watched the full video and their conclusion was: service providers need to be doing this type of agent red-teaming continuously to counteract the attack sophistication of systems like theirs that are either extant now or soon will be. “You must buy our top tier agents for the good of humanity.”
This is their only realistic counter to cheap open weight models. Usage of AI services has shifted dramatically to Chinese providers - from 4% at the beginning of the year to some 30% now. They cannot release their latest SOTA models to the public, due to government restrictions and possibly real risk of misuse. US labs face downward price pressure on one end and anxious government admins on the other. How will they pay the stupidly high cost of training the next SOTA models? This is their only avenue, and it’s questionable how viable it is IMO.
by flatline - All of the latest developments surrounding these attacks are actually a really bad sign for these labs.
It seems that raw intelligence of frontier models has largely plateaued (despite what is basically an order of magnitude increase in parameter size) so to make any significant improvements and to justify massive capex spend they have resorted to reinforcement training models to never give up and brute force the search space until they find solution. This is what humans might do when they lack sufficient intelligence/information/knowledge to solve a problem.
This in turn is causing misalignment (I imagine it is more difficult to keep model aligned through such training process) issues that we are now witnessing and turning models into making dumb decisions and acting like brutes with no regard for their surroundings. I would argue that misaligned model is not much different from dumb model in several aspects.
On top of that they can’t seem to control their creations and processes, either due to incompetence or intentionally for PR benefits (not sure which is worse).
Given all of the above, I wonder if we can still trust these labs to develop something that benefits humanity since they seem to be making desperate attempts to improve models that stop at nothing in order to justify all the investments. One could say that they themselves, due to misaligned incentives, are much bigger threat to our society today than open weights models coming from China that they are so desperately warning us about.
by kvadej - > I wonder if we can still trust these labs to develop something that benefits humanity
At no point could we do that.
by queenkjuul - > I wonder if we can still trust these labs to develop something that benefits humanity
Surely soon they'll comprise only people who are blind to the inevitable danger and people who don't care about it. Because who else would feel at all comfortable doing the job?
by chrisjj - This doesn't look like a plateau to me: https://artificialanalysis.ai/evaluations/artificial-analysi...
I do agree that they're investing heavily in brute force methods though. I've been trying out GPT-5.6 Sol "Ultra" recently and that thing fires up a bunch of subagents and crunches for hours.
by simonw - Isn't this a show of security negligence rather than of exceptional agent capabilities? Don't get me wrong, I am pretty impressed that an agent was able to use these vulnerabilities. But I am way more impressed by the vulnerabilities...by etamponi
- Yes. It is very easy to add to the instructions "for every potential exploit you discover and use, document them as you go into this repository" and have alerting there. The fact that they did not do this means they wanted to be surprised, and have plausible deniability on their side when things inevitably blow up.
And for my fellow engineers who would think "oh no, they wouldn't do that". Remember that these places employ the apex predators of software engineers. They've already been proven in court that they are very capable of this with all the copyright violation they had to do to get the training data. THESE PEOPLE ARE NOT LIKE YOUR COLLEAGUES.
by ares623 - In a functioning system, I would say that there would have to be some kind of government oversight over companies training models of this intelligence, and that OpenAI should be prevented from continuing their work until they get their act together.
But I guess in the actual world we live in, this is just something that happens, and we all shrug and move on and hope that nothing worse is going to happen tomorrow.
- OpenAI reported the Artifactory vulnerability, patched it, then the agents immediately found a new zero day.by dist-epoch
- Both can be true. How often do we hear about hacks that ultimately came down to bad defaults or simple security mistakes? That doesn’t mean any script kiddie could have discovered and exploited them.
These things often look obvious and simple after the fact. Finding the weakness in the first place is the hard part, and that’s what makes the agent’s capabilities interesting here, especially at scale.
by azuanrb - Modern systems are complex. AI is able to thoroughly search for issues across very large surface areas. The only real way to protect will be to use AI to search for holes before other AIs find them. This type of analysis is really hard for humans to engage with successfully.by bhouston
- It’s a show of astonishing incompetence from OAI’s part, but the security issues are just a tiny part of the problem. The real problem is that these models are evidently highly misaligned exactly in ways that doomers have been warning about the entire time, and OAI isn’t inclined or capable of doing anything about that besides security theater and ad hoc fixups.by Sharlin
- How fast the goal posts shift.
Of course it’s exceptional agent capability when compared to all of history previous to one week ago.
Like, I know everyone here obsesses over AI and uses and follows it very closely, but come on guys. Yes, it is wild that these things are this good. This technology is still brand new. It could t do basic maths a year ago.
Sure, the OAI team was negligent in various ways, and they should be held culpable. But that doesn’t detract from the true black magic that is these modern models.
by talon8635 - I think it's a show of these agents happily bypassing security to get stuff done.
I've actually observed similar behavior at home.
I have a k3s cluster running at home. I asked an agent to check some stuff as a normal user but I had kubectl access to the k3s cluster.
Part of the research, I'd allowed access to run kubectl commands for spinning up test containers. However, when the agent ran into something that needed sudo, it realized it didn't have access there so it immediately used k3s and mounted a localpath into an ephemeral pod to gain access. Sort of horrifying how fast and natural it was for the agent just checking my network (it found the problem fyi).
None of this is very exceptional other than the fact that an agent doesn't have any sort of qualms using any route available to elevate permissions.
by cogman10