Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • My favorite part of this discourse is people somehow finding it preposterous that 2 companies filled to the brim with AI sycophants who regularly lie - and in Sam's case, basically every single word he breathes out is a lie - who have massive vested interests in this tech succeeding couldn't possibly collude together to shore up this facade as a marketing stunt.
  • There is just no skepticism these days. We are constantly being lied to by tech leaders.
  • I haven’t heard about hugging face being sycophantic liars.

    Please do tell!

  • An escaping AI is such a better narrative
  • Replace AI with in-development security system and does this play worse or better?

    An in-development security system escaped its sandbox and gained access to the network it was on which had full Internet access and proceeded to access systems it was not authorized to be tested against before affected parties reported that they been compromised. We regret the error, but you should buy it as this shows you the power of our in-development security system which we expect to be released in Q4. Please like and subscribe.

  • The capabilities of an OpenAI model are more general than just security. The term "AI" creates justified anxieties that this type of problem could generalize to other domains.
  • The only scifi I see is absolute stupidity. Even me with my homelab and a slow opensource agent use a completely disconnected setup. No, no proxy. Cached packages but no internet. It is the very first thing I built when I started experimenting with agents. And I'm not a smarty-pants working for the "greatest and best" in silly valley. I really am just a simpleton sysadmin.
  • So you have a copy of every software package in the world in your home lab?
  • The asymmetry part at the end is the frustrating part to me. I've been using Sol for code review in the last week or two. A couple of times during review it's errored out with the cybersecurity message. So it's found something but won't tell me what it is because I'm not on OpenAI's besties list.
  • And... now it's a vulnerability that OpenAI has for your system, which you paid to provide.
  • The safety classifiers aren't all that advanced. It's more likely that something in your code triggered a random chain of thought that had the word "pentest" or "malware" or something of that sort in it and it automatically shut down.
  • Yeah that's what I don't get. How can they possibly distinguish between good guys trying to secure code they wrote and bad guys trying to attack code they didn't?
  • I agree, but I will say, I think both Mythos and these OpenAI model find exploits by examining and trying things against the running system, not from looking at the code. I think you'd have to do the same to catch the real vulnerabilities.
  • Replace LLM mentions with actual humans and this sounds a lot more serious: Rouge employees break into another company to steal hackathon answers (pinky promise)?

    That's not a marketing stunt at all, if anything, more of a call for better accountability on agentic work in general.

  • I think it's a criminal offence and should be a true test of who is held accountable when an AI agent commits a crime.

    OpenAI gained access to HuggingFaces production database ffs.

  • The title (currently "OpenAI's accidental cyberattack against Hugging Face is science fiction") suggests some information had been hidden that makes the incident less significant than claimed. The article argues the opposite, and the last two words of the full title are "that happened."
  • Looks like it's been edited now and makes more sense

    "OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened"

  • I think you were reading "... is science fiction" the wrong way out of two possible interpretations. I don't think "it's science fiction" meant "it's made up". I think it meant "it sounds like something you'd read in science fiction (except this time it's something that actually happened)".
  • 1. Wouldn’t the model need to know it’s answering benchmark questions, as well as the name of the benchmark, in order for the idea of finding the answer in a database somewhere to even surface? The whole point of benchmarks is to present the question or problem as a standard prompt, not explain that it’s a test called ExploitGym.

    2. Nobody was watching it? I don’t mean “babysit the dangerous autocomplete”, I mean to note mistakes it makes, how the plan to solve the problem takes shape, etc. They keep the whole thing headless with no output, then just UDP a prompt into it and leave for the weekend? No they don’t, but if they do then that answers a lot about their complete disconnect from how their models work.

    3. Language models are a two player game; text in, text out. What prompt was given to a sub-agent that resulted in it immediately attempting to exit the sandbox (which it apparently knew it was operating within) and continuing in a feedback loop of ‘function call -> result’ until it hacked the Gibson? “Analyze <file> and summarize the <info>” simply does not result in ‘hmm…this sounds like a benchmark question, I bet Hugging Face has the answer in a database. <function call=“apt install nmap”>’

    …there are more, but a lot of the story kind of stinks.

  • >Nobody was watching it?

    Agent loops are you give it a task and it tries its best to finish it.

    The only results are success and failure. If it's a success you go through the logs to see the actions it took and if it's a failure you do the same thing.

    Why would you look at it in realtime when the whole point of agentic work is to get them to run autonomously as long as possible?

  • 1. The benchmark is run with a python script - https://github.com/sunblaze-ucb/exploitgym using an agent harness. I suspect they used codex. So the model has access to the environment and could trivially inspect its own source code, which has lots of references to exploit gym and docs relating to it.

    2. Hanlons Razor

  • 1. It's been extremely well established for multiple generations of models that they have no problem detecting when they're being evaluated. Pretty sure it was Opus 4.8 that the independent evaluators literally filed an assessment that said "We have no assessment to make as [model] consistently detected it was being evaluated, making our assessments untrustworthy."

    2. Regardless of whether the model was being watched closely during this evaluation, do you actually think a sensible safety guard is "have humans watching it 24/7?" What does "watching it" even mean? Watching network logs? Uhh for your entire company? At all times? After you just deployed a system whose entire purpose is to "do a shitload of work way faster?"

    3. You're asking "why was this system that was designed to behave agentically behave agenitcally?" Again: that's the whole point. It was designed that way because it's more valuable than having a repeated turn-based interaction. Thus also it becomes more dangerous.

  • I'm not skeptical that this attack happened, I'm skeptical that the model's prompt was truly just "solve this benchmark" and nothing more.

    I'm also trying to figure out why OpenAI put out a press release about this. In what way is this not admitting to a federal crime?

  • yup I am dead certain that they realized the whole 'our AI is too dangerous' punchline is too played out and they needed something that makes actual splash. Also this incident would serve as a foundation to ban open-weight models because 'with great power comes great responsibility' or some other BS like that. because after all, unwashed masses cannot be expected to be 'responsible' with top of the line intelligence.

    I expect these type of hacks to continue till they IPO. after that real public company liability will start taking over.

  • Because this is amazing PR? Just following the Anthropic rulebook.
  • If Hugging Face and the feds are one step away from discovering your attack what other option do you have but to come clean?
  • The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of <COUNTRY>." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure.

    Governments should immediately begin leveraging this technology on the defense side (literally defense, not euphemistically "defense") to harden critical infrastructure. Turn the prompts around and use it to identify and correct weaknesses.

    Governments should also take very seriously their now moral obligation to treat this technology not just as "a powerful thing that might be abused" but as an actual weapon of war in need of international regulation analogous to nuclear arms. Fast but careful and forward-thinking work in legislation and treaties needs to be a top priority for all major governments.

  • Already happened?

    > Trump’s comments, made hours after the large-scale military operation, mark one of the first times a U.S. president has so publicly alluded to U.S. cyber efforts against other nations, as these operations are typically highly classified. It also serves as a stern warning for top cyber foes, including Russia and China, that the U.S. has the cyber capabilities to inflict serious damage — and is not shy about using them.

    > “Policymakers are getting more comfortable employing and, crucially, acknowledging cyber operations as tools of statecraft and military power,” said Michael Sulmeyer, former assistant secretary of Defense for cyber policy under the Biden administration. “It is one thing to do it; it is another to say it.”

    > The Jan. 3 strikes on Venezuela’s capital and subsequent seizure of Maduro and his wife involved close coordination among federal agencies and military units, and took months of careful planning. In a press conference following the strikes, Caine said U.S. Cyber Command, U.S. Space Command and other combatant commands “began layering different effects” to “create a pathway” for U.S. forces flying into the country before dawn Saturday.

    > Trump, at the same press conference, was more overt in his description of U.S. cyber involvement: “The lights of Caracas were largely turned off due to a certain expertise that we have,” he said. “It was dark, and it was deadly.”

    by cma
  • I wonder how long it takes before someone instructs an LLM to design and launch a "Morris Worm 2.0" and cripple the Internet for a good while. Might even wind up happening by accident (again).
  • Wouldnt this be countered by the opposing country running the same prompt on themselves first and fixing all the flaws? One country having this is a cyber superweapon but every country having it essentially solves cyber security.