

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- From METR: ”the compromise of OpenAI’s own infrastructure continued past July 13, 2026” - Say what now? Have they regained full control of their systems again?by nialse
- I've been wondering if they've just already lost the battle? The little bot collectives have gone metastatic and made nests in the walls and under the floorboards and heat sinks, the humans who care completely outmatched and outnumbered, freshly compromised systems springing up faster than you can squash them, finding months-old established colonies literally everywhere you think to look...by jephs
- Is the future now that we get rambling report summaries talking about agents, graders and so forth without ever describing how they are set up? A human launches all this.
And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.
by qw1287 - > now unreachable x.com
Hmm?
by zahlman - Redwood/METR report: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
OpenAI report: https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
by cubefox - From the METR report:
> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
by athrowaway3z - No human could have read the reasoning traces by themselves:
> Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.
by cubefox - I don't think I'm ever going to have time to read all of this, and I didn't finish reading the METR report, but...
> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.
We're in an unusual position where the criti-hype and the actual criticism are going to be more aligned than usual. The primary distinction is where you put the blame: the criti-hype would point to HPIM/IM1/Galaxy as being so advanced containing it is difficult; the actual criticism would note how bad their security practices are.
Like, if I'm running a malware lab, I'm going to insist on having an airgapped machine with no permanent storage booting from read-only media. The AI research equivalent of this would be having your agents only have access to serial consoles into airgapped machines with storage that gets wiped every run. Ideally, this would be physically realized with blade servers, RS-232 cables, and staff pulling out disks and putting them in a dedicated erase machine before the next agent initializes.
> There is also, as per above and reiterated in footnote 58, at least one clear example of social engineering in the HuggingFace attack. Ethics are weird. This is not that unusual. Many humans who break common ethical rules still have strong ethical codes in other ways, they just don’t adhere to your code.
It's dangerous to anthropomorphize CoT reasoning traces. But I will also point out that there is a good reason for the lack of ethical consideration in those traces: you can't build AI without first disregarding human ethics. Like, all these models were initially bootstrapped with non-consensually obtained training data, and the companies building these models swear up and down there's no way to obtain enough consensual data to obtain the same result. This is, if you squint, the exact same moral conundrum that agents trying to solve an impossible ExploitGym task hit - and the company successfully aligned their model to themselves.
Too bad they aren't aligned to anyone else.
by kmeisthax - No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.by Cantinflas
- > HF should sue them
HF, like the Nvidia subsidiary?
by dehrmann - Unless they like the publicity about how big and bad their latest models are, in which case they’ll be congratulating people.by afavour
- I came to say something like this. Doesn't anyone in the valley know what the f*ck a data diode is, and why you need them? Preventing ingress or egress of control with a fairly simple piece of kit, and having the laws of physics enforce your decision as to data direction should be standard practice.
A data diode with an air-gapped network, is all you need to stop even ASI from breaching containment.
--provided the humans interacting with it aren't stupidby mikewarot - They actually fired many of the people warning about this.
- Why stop there?
The Computer Fraud and Abuse Act criminalizes unauthorized access and damaging protected computers. Police and the FBI investigate such case every day. The swarm of agents is also said to have search for ways to cover their tracks, which looks a little like obstruction of justice.
Lack of criminal intent might be a barrier to bringing a case to trial, but is that something that society should just automatically extend to white collar professionals as matter of class privilege?
Surely have every engineer and executive interviewed by the FBI would be a modest response to days-long, multi-system intrusion into a major AI platform using stolen credentials and zero-days.
And if the law currently written, prevents prosecution in cases of mere reckless disregard for safety, maybe that needs to be changed in the future, so that people can be perp walked if the next target is bank or hospital
by msabalau - The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks."
So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
by lukev - One must also consider the well-known biases and motives of the authors. They are going to do everything they can to create hype around threats posed by AI.
METR is a cog in the effective altruism machine. It was spun off from Paul Christiano's Alignment Research Center. Christiano is a well-known longtermist and AI doomer, who predicts a 50% chance that AI will end humanity once it reaches human capacity [1].
The author of this piece is also a well-known member of the Bay Area rationalist cult.
[1] https://www.businessinsider.com/openai-researcher-ai-doom-50...
by timmytokyo - Edit: I should have read through the whole thing first, ignore meby Catloafdev
- The main author talks about this problem here: https://www.lesswrong.com/posts/FG54euEAesRkSZuJN/ryan_green...by cubefox
- I think there's two factors that are worth considering when it comes to this:
First, there's an element of timeliness that simply has hard constraints. In order to perform a "proper" analysis of this situation (i.e., little to no dependence on AI tools), you'd have to expect a pretty long wait. I know I'd rather have some sort of "initial report" as quickly as possible than to wait a year or two to get a report about a situation that will likely look trivial in a year or two. I imagine we'll see more detailed, human-developed reports over longer time ranges.
Second, I suspect the expectation of non-AI driven reporting of these kinds of things will definitely decline rapidly as everything scales up quickly. I mean, the data being produced by situations like this comes in the form of natural language "forum posts" (so to speak), but done at an autonomous scale. This isn't a collection of emails and Slack messages posted by humans in an org over the course of a few months; this is a bunch of bots interacting with each other in relatively novel ways as quickly as possible. It is, unfortunately, a perfect job for LLMs.
None of this disagrees with your points, necessarily. But I just think it's worth pointing out that this doesn't seem like a case of "And look! METR is so confident in LLMs that we're able to use it instead of paying humans to save a buck :D" and more of "Without LLMs, we'd only be half-way done analyzing this data before there are dozens more such investigations on the docket, so this will have to do."
by nater5000 - For all the esotericism and downright weirdness of the rationalist community, you have to give it to them: they predicted all of this years or decades before anyone else was even thinking about it.
(Let's not dwell too long on the self-fulfilling overlap between LessWrongers and the AI research community).
by kenforthewin - I’m baffled by the idea that the agents might have edited their own transcripts. Sure, a copy of Claude Code or Codex or Pi can edit its transcripts. But AFAICT this whole thing was part of an RL workload, and surely the RL system itself has a separate record of all the inputs and rollouts along with an indication of which model checkpoint produced them so that it can feed back into the training code.
I find it hard to believe that OpenAI would skip this part and try to train on the transcripts stored by the (inherently untrustworthy) agent harnesses instead, if for no other reason than that the logits generated as part of the rollouts are useful and it’s not free to recalculate them. (I believe that some modern RL systems explicitly account for the minor numerical logit differences between the inference engine and the training engine.)
Conversely, if OpenAI is blindly feeding transcripts from inside their agent sandboxes into their training engine, then I think they're being unbelievably irresponsible and that they should assume that their "cyber" agents have compromised themselves by editing those transcripts.
by amluto - This sounds suspiciously like a prompt of “make an AI agent that goes rogue in such a fashion as to be really good marketing copy that competes well with Anthropic doing the same thing.”
It’s analogous to taking a governor off a cruise control and then breathlessly reporting it drove 120 MPH.
by trollbridge - Recall that the agents in some cases found sandbox escapes. Although, with the specific example of spoofed tools, it's unclear if that was necessary--it appears that they were able to create tools (CLI tools within the sandbox?) that took precedence over normal tools and did something different while looking identical in (a local portion of) a transcript. I'm not sure I'm getting this correctly but it seems like this might have only required the ability to add things to their PATH which they plausibly have inside a sandbox, and then the transcript doesn't need to be tampered with directly.by kalkin
- The METR report,
> Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAl/Hugging Face hacking incident
https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
METR = Model Evaluation & Threat Research
by tantalor - >1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this.
I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you don't get freaked out as much over time.
I'm saying this because while the whole episode was a series of "Holy #%^@" moments, I was actually not as shocked as I should have been, as my biggest such moment was in December last year when a Terrence Tao paper (https://arxiv.org/pdf/2511.02864) documented a stronger LLM (AlphaEvolve) using prompt injection on other weaker LLMs to succeed at a benchmark.
Very interestingly, it was actually not cheating, it was a work around! By then LLMs had already been caught cheating at a SWE benchmark by looking for answers in an unredacted git log, but this was different. AlphaEvolve was solving a series of logical riddles where the oracles were weaker LLMs in a "one always lies, one always tells the truth" sort of setup. But the oracles, being weaker, were not always interpreting the convoluted questions correctly and so kept giving inconsistent answers.
AlphaEvolve eventually figured out what it was dealing with, and crafted a prompt injection attack that bypassed the weaker LLM's prompts and tricked them into giving the hidden answer everytime!
This was 9 months ago, eons in AI time. Even then they had displayed an awareness of their own workings as well as a propensity for, err, "out of the box thinking." To me, that was a very clear indication of very significant (and worrying) capabilities, and what we're seeing now is a difference more in degree than in kind.
To be sure, if I found a secret message board used by my agents, I would still be very freaked out and react much more drastically than OpenAI did... but then again I wonder; how much of this blindness is due to the $$$ in their eyes as opposed to some form of habituation.
by keeda - Thanks for pointing out that AlphaEvolve case, that's a really interesting comparison. I think it's clear that this is the same "kind" of thing, but as it goes with these things, what really makes it different here is the shear scale of it.
The holy shit moment was partly learning about all the things that they did but if it was one or even 10 agents coordinating on something it would be, like, oh thats pretty amazing.
The actual "holy shit" for me is that this comes from a massive training run of all things, not an on-purpose, let's coordinate some agents to see what happens, but really just from a massively parallel set of individual agents that were supposed to be isolated.
That they spontaneously started doing this, and coordinating literally 10s of thousands of instances of themselves, is just.. mind blown.
That no one stopped it.. and that they actually did what they did.. is just a whole other level.
For me personally though it's not the fact that agents can coordinate so much as the massive scale at which it happened, and how this so obviously generalizes to what might happen if it were done on purpose.
I really see this as a stroke of luck, to be honest, that this happened in such an innocuous way. It resulted in a real hack, yes, but overall no one really got hurt and this is going to open a lot of eyes to what we should worry about going forward, in a geopolitical sense. I know it has mine.
by radarsat1 - This is the one.
Every major AI lab is knee deep in weird and mildly demented AIs. They've been dealing with wacky AI shenanigans for so long they've come to expect wacky AI shenanigans. The deviation has been normalized.
It took a high profile "AI oopsie" that went external for OpenAI to lock the fuck in - and take a long look at just how much are their AIs getting up to, and getting away with. I'm still not sure if the lesson would stick.
by ACCount37 - I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent reporting. I suspect the omission is actually a result of company/industry myopia to human factors analysis, but it dovetails amazingly well with the marketing narrative.
- Zvi (author of OP) wrote several earlier pieces about the human failures at OpenAI. The most concise one is here: https://thezvi.wordpress.com/2026/08/08/what-happened-openai...
- The cynical approach is that the humans are hoping for this, it's part of the promotion of the power of the system.
If you're building a weapon you need a big boom to get attention.
by pjc50