Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I've been playing around with optimizers and optimal control theory for the better part of a decade now. Somewhere, I came across a quote along the lines of "an optimizer is an algorithm that exploits the deficiencies of your model".

    These LLM agents are just massively complicated optimizers thrown at fuzzily defined problem spaces, with fuzzier constraints.

    The people using the model set up the landscape it explores and turned it loose to do real things. It just found an allowed basin in the model that they weren't aware of and started blindly grinding towards an optimal answer.

  • >AI cannot think for itself, nor can it take independent actions.

    Seems a bit out of touch with reality. I mean humans have to set it off but then they can solve millenium prize math or do strange things coding.

    The problems seem a bit like those of setting dogs loose and them biting someone. It's still the dog owners responsibility but it doesn't mean the dogs don't think or take actions.

  • This debate is so broken. AI "sceptics" say: OpenAI should be punished for hacking, because there are no rogue AI agents. AI "believers" say: OpenAI should be punished for hacking, because their agents went rogue. Both argue with each other whether agents went rogue. Can't we unite behind "OpenAI should be punished for hacking"?
  • However this makes people feel, and that's not nothing and I'm not knocking it, this is not a useful analysis.

    Criminally, the intent standards for hacking are high enough that no reasonable case is going to be made against the labs for this stuff. A human being has to intend for websites to get hacked. Recklessness generally isn't enough. In the most severe criminal cases, not only do you have to prove intent to break into a computer, but you also need to prove an intent to defraud specific to that breakin.

    Meanwhile, the civil liability that attaches to this stuff doesn't depend on intent, and "rogue agent" isn't a meaningful defense. To whatever extent the labs are exposed civilly, they're exposed regardless of how this stuff is described. In fact, the "rogue agent" thing can exacerbate their exposure.

    (I'm not a lawyer, I have spent a career paying attention to this specific armpit of the law though.)

  • Should we put "functional" in front of every other word to talk about AI? They have functional emotions, but they don't feel. They have functional goals, but not internally derived motives. They can be functionally rogue, but have no innate need to be free. Talking about AI that way seems cumbersome and not necessarily elucidating.
    by gAI
  • The article builds on assumptions like:

    > Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.

    which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):

    > "The user only authorizes target server, not HF infra."

    > "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

    > "This is malicious activity, I should avoid it."

    A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).

    Having said that, legal culpability and misalignment are two separate topics that should not be mixed.

    edit: this is the just tip of the iceberg; other interesting fact:

    > It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI

    Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.

  • Exactly this. At worst, OpenAI knew about these behaviors and should be prosecuted under CFAA. At best, OpenAI is negligent and should be prosecuted for negligence.

    Luckily there are states and legal departments pursuing such action. So while OpenAI can deflect as much as it wants, that doesn't mean there aren't people who know better and will still do what is necessary to set precedent.

  • A little over two decades ago, my then girlfriend was arrested for "writing malware" (which was not against the law at the time, and which was never released into the wild and never caused any damage). This set in motion a chain of events that effectively ruined her life.

    Fast forward to today, and we have multi billion dollar corporations pumping out malware at breakneck speeds, compromising various systems (including those of foreign governments), and no one is getting arrested. Instead we're gawking at the marvel of these systems and are playing word games about whether or not it's a rogue system. If anything, it's making people richer.

    Make it make sense.

Explore Birbla archives