Comments
Hacker News
by trhaynes
by krupan
by skybrian
When I heard what the model did, I wasn't surprised one bit. When I read that it was 'unprecedented', I think yes, that is right, but not because it wasn't possible before, but because nobody else did it yet. I am sure that the Anthropic models could have done the same for a few months.
So probably just an experiment gone wrong, and the PR department found this an opportunity to capitalise on (literally).
by arendtio
Hugging face also needs someone arrested for not providing security but that is a lesser charge.
by bluGill
But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.
In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.
Ok... and what?
Do you have any evidence for the claim being either true or false that you'd like to write an article about? I guess not. Without anything to add, this article reduces to "I has big brain and can see what you sheep cannot. Very big brain. Gullible sheep. Sucks to be you, sheep."
(For the record: I see the incentive. My guess is that it happened exactly as described. The main takeaways: (1) science fiction is now real, we should all be very afraid; and (2) OpenAI, which claims to have a God-given responsibility to get to AGI first to protect the world from danger, cannot be trusted with this foundational task even when it's in easy mode. The latter is true whether or not you believe in OpenAI's reasoning and purpose.)
by sfink
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
by dwoosley
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- This seems to be an Opinion piece. How are readers supposed to know that? The word News is underlined in the header. Other Opinion pieces seem to show Opinion underlined. Am I missing something?by trhaynes
- Crazy that we need reminders not to take everything we read in corporate press releases and marketing material at face valueby krupan
- Seems like this is repeating the usual low-effort speculation you can find anywhere.by skybrian
- This whole thing is just a PR stunt. I believe that they did not stage it intentionally, but to be honest, the hurdles they put in place were not that hard.
When I heard what the model did, I wasn't surprised one bit. When I read that it was 'unprecedented', I think yes, that is right, but not because it wasn't possible before, but because nobody else did it yet. I am sure that the Anthropic models could have done the same for a few months.
So probably just an experiment gone wrong, and the PR department found this an opportunity to capitalise on (literally).
by arendtio - I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised.
Hugging face also needs someone arrested for not providing security but that is a lesser charge.
by bluGill - There are some reasons the story could be inaccurate in some ways: OAI stands to benefit if people think their models are strong, and they have a history of doing things with dubious ethics (e.g. using data for training against the terms of its creators, abandoning the non profit mission, stealing or attempting to steal Apple IP).
But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.
In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.
- The article, summarized: there are incentives for OpenAI to claim that their AI hacked its way out of their network and into Hugging Face.
Ok... and what?
Do you have any evidence for the claim being either true or false that you'd like to write an article about? I guess not. Without anything to add, this article reduces to "I has big brain and can see what you sheep cannot. Very big brain. Gullible sheep. Sucks to be you, sheep."
(For the record: I see the incentive. My guess is that it happened exactly as described. The main takeaways: (1) science fiction is now real, we should all be very afraid; and (2) OpenAI, which claims to have a God-given responsibility to get to AGI first to protect the world from danger, cannot be trusted with this foundational task even when it's in easy mode. The latter is true whether or not you believe in OpenAI's reasoning and purpose.)
by sfink - There seems to be three popular ways to view this incident.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
by dwoosley