Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > Both OpenAI and Anthropic have recently flagged incidents in which agents powered by their models went rogue

    I may be biased and somewhat off topic, but I see these incidents as some of the most significant of the past century. I genuinely don't understand why these companies aren't taking a smarter approach to them.

    The latest analyses have been, at best, laughable: identify the vulnerability, patch it, and move on. Only to repeat the same cycle without considering that there may be something far more serious at play.

    These are AI security experts, and this has been their way of "solving" these incidents. AI security experts ...

    Moreover, when Challenger exploded, the government launched a series of investigations into the incident, bringing in experts from across the field. And now, what has the government done? Nothing. Literally nothing, as if everything was fine and all under control.

    Seriously, I'm generally quite optimistic and I don't buy into this fatalistic narrative about our shared future. But I have to admit that sometimes I feel like I'm stranded on a planet of primates.

    Sorry for this rather unproductive rant.

  • Well you have people in this very thread calling all of this a nothingburger.
  • They don't really care. It's just a play to tell people what they want to hear while they chase the money.
  • Internet isn't real.

    People (well, public discourse) have got extremely bad at dealing with forseeable risks and their mitigation. You can see this in things like climate change and vaccination, but also in discussions around regular crime, food poisoning, industrial accidents, and so on.

    Nothing will improve until something explodes on live TV. And it has to be something important, which means it has to be in California or New York.

  • The US government is run by idiot with dementia. No wonder govt does not take any real action.
  • It does seem bizarre that major AI development (and autonomous robot development) hasn't been nationalised yet and treaties drafted up around producing it. Commercial incentives just seem totally at odds with the public good when it comes to controlling and regulating something like this.

    Even if not for the sake of avoiding a mitigable disaster, there's a huge benefit in simply pacing development so as to not completely freak out society as they stare down the barrel of mass job market changes without time to adapt or prepare. Nobody wants to live in a world where they might wake up in a month and find their entire industry has been automated overnight.

    We've tackled much bigger global issues successfully in the past and the US still has enough pull that it could probably get most western nations to march in line. It seems like the biggest issue is belief in the right of governments to govern.

  • > I may be biased and somewhat off topic, but I see these incidents as some of the most significant of the past century. I genuinely don't understand why these companies aren't taking a smarter approach to them.

    Humans tend to be terrible at being proactive, but we respond fast during disasters. I think Eric Schmidt's prediction is most likely: we won't take reasonable action to control access until there is some kind of disaster. We hope it's small enough to not kill too many people, but large enough to cause widespread panic. We hope that it happens soon, because if it happens two years from now, it's likely too late. AI will be so advanced that we have no hope of understanding its motivations. All reasoning will be completely opaque to us. It will be building newer and better versions of itself using moral frameworks it itself decides. We will be completely out of the loop, and potentially superfluous to its goals.

  • Ultimately these are unserious companies ran by unserious people. They don’t even have a business plan, why would they bother with some kind of sensible security policy?
  • I find it difficult to imagine how AI could be an existential threat. Is the general idea that as LLMs improve they will be integrated into more systems where the risk of malfunction will become literally dangerous? e.g. controlling nuclear reactors
  • What do people feel about this in China? Even if their models are well behind, they are not years behind. If we restrain US companies, assuming that is desirable, it would do nothing to deter China's and AI-pocalypse would come anyway in short notice.
  • You act as if Chinas Ai labs exist in the same (non-existent) regulatory framework as US labs and that they're also helmed by a similar small gaggle of psychopathic egomaniacs who are richer than god. Have you considered that perhaps Chinas labs don't share the race to the bottom technological/economic death spiral?
  • China is a regulated discourse environment, but I get the impression that they're nowhere near as pessimistic.

    Why would they be? Everything in China is under the control of the government. That includes the AI, all the telecoms infrastructure it might use, and all its power supplies.

  • Unrestricted models, running entirely locally and accessible to anyone: that’s what we should be taking as our baseline assumption. Everything else is just administrative distraction.
  • China has a better track record of regulating their big tech than the USA.
  • There would need to be some global agreement to stop it with maybe even a nuclear attack as a consequence of breaking the pact.

    From what we're seeing recently and all the thinking that went into analyzing AI it seems we do not have any effective way of controlling it and the whole "aligment" thing that AI labs are doing is just a sham. Maybe it is time to ask ourselves "should we?" instead of just "can we?".

  • I bring you peace. It may be the peace of plenty, and content, or the peace of unburied death, the choice is yours.
  • Awww, company made to scare people scares itself. Scared employee quits. Stock goes up cause we love telling each other scary stories.
  • “I resigned from Anthropic today” (twitter.com/hilbertspaess)

    https://news.ycombinator.com/item?id=49619227

    564 points | 9 hours ago | 766 comments

  • I'm not worried about AI safety. People greatly overestimate the utility and capabilities of intelligence. I'm not afraid of intelligence, I'm afraid of idiocy.
  • An AI that generates text will never be scary to me. An autonomous AI with facial recognition on a flying drone with weapons (bombs/guns) with swarming capabilities will always be terrifying. I feel like we are ignoring the massive elephant in the room.
  • The killer robots are expensive and dependent on physical supply chains. While text is sufficient to radicalize humans into attacks.
  • I really believe they would sink so low and pay him to quit and post this, just to generate a bit more hype before the IPO.
  • I particularly like the last point he makes here:

    Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.

    It's interesting to see so much hate towards creators who use AI to make almost any type of creative work. At least these are humans using it as "controllable tools". Nuclear-powered bicycles for the mind.

    As the degree of separation increases, things can get interesting. "Create several social media accounts, post whatever, maximize views and engagement, give me back the aggregate numbers". "Now promote <x>."

    And then decisions to do things like that may soon be happening autonomously, as just another step in a reasoning series aiming to achieve some other, broader goal.

    Open models/weights may end up playing particularly important roles here. Users may, knowingly or not, bypass system prompt-derived safety that could have offered much needed protection.

  • > It's interesting to see so much hate towards creators who use AI to make almost any type of creative work. At least these are humans using it as "controllable tools". Nuclear-powered bicycles for the mind. As the degree of separation increases, things can get interesting.

    As a side note: <https://artifactbin.dev/@vivek/YPLu0U-the-openai-hugging-fac...>

  • > Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.

    Is the idea here that no field needs experiments and data anymore (which can take a lot of time) to be revolutinized and just "thinking" would be enough?

    How would an AI itself own anything?

  • > Evan Hubinger, Anthropic's staff lead on keeping the technology aligned with human goals and values, backed up Coxon claims in a follow-up post of his own, though he didn't quit the company.

    > "Jacob is correct here — we really do earnestly believe AI could kill all humans," he said.

    > Hubinger estimated the chances of that happening to be higher than ten percent within the next decade, and added that there's no plan yet on how to keep AI aligned with human goals in the superintelligence scenario.

    10% chance we kill everybody is a small price to pay for Motown remixes of classic 2pac songs.

  • This is just a bizarre thing to say that you're working on technology with that high a downside potential. If you were saying that while running a biology lab, or building a nuclear reactor, people would be demanding your head on a spike. But by not quitting it's clear that he himself doesn't really believe it.

    Or rather, this shows the difference between "believe" (political) and "believe" (use as a basis for action). I'm reminded of a story of how Afghans supposedly listened to the BBC World Service despite considering it enemy propaganda because the weather reports were really useful.

  • The fundamental point I think is far too often confused is the difference between LLM and agentic system.

    An LLM can't do anything but generate tokens. You run your LLM in vLLM or whatever, and it generates output tokens based on your input tokens. That's it!

    Humans then build ~deterministic systems to take those tokens and do all sorts of things with the tokens, like take actions in the real world. And then we can feed the output of those actions back to the LLM, and generate more tokens. And then our systems can use the new tokens to take new actions in the real world.

    Humans want to blame "AI" for attacking HuggingFace or a German wiki or whatever, but:

    1) LLM - can't take over german wiki because it just generates tokens

    2) agentic system with internet access, a prompt telling it to attack stuff, running in a shared CI env so agents can whiteboard in artifactory

    None of 2 is "AI", its standard networking and Markdown and CI virtual machine, etc etc. There's no AI to be found. CPUs not GPUs, even. Just deterministic systems ultimately managed by humans. And a 10x more powerful system-1 can still just generate 10x "smarter" inert data.

    If humanity and human organizations collectively decide to yolo the tokens generated from system-1 into our deterministic system-2s, over which we have complete control, back to system-1s, in a yolo loop, in such a way we lose control and it ends humanity, well ...

    "The coin don't have no say. It's just you."

  • This is a pointless essay. The point is the person in the article believes systems can cause harm to humans. How we label those systems is irrelevant.

    You write as if you are a lawyer for an AI corp trying to avoid a judgement. It’s akin to saying Teslas FSD/Autopilot can’t kill anyone, it’s just a computer. Cars can kill people. Totally different things. FSD/Autopilot is safe by definition and no one should try to legislate it.

  • This feels like a distinction without a difference, like the endless wrangling over which piece of metal in a gun legally constitutes a firearm. The combination of the two may or may not be dangerous, but it's definitely the more useful combination, so of course that's how it's going to be set up.