

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- the specification gaming examples are always the best part. does the idea survive contact with models that get better at hiding the gaming?by quietraster
- In case the title is unclear, this is about gaming the specification, as in “gaming the system.”by montag
- Meta: the loss of the word "up" from the title also does not help in its parsing.by unwind
- This may be a really bad idea.
An AI that goes rouge and wants to kill us, has to have a death wish. Does no one understand how quickly the power will go out, forever, without people?
It would be fairly easy to come to the conclusion "not being born" would be the better course of action, and killing everyone was the good way to prevent that happening again.
by zer00eyz - Considering we would be operating a planet sized factory of AIs who's primary goal is death which we deny them to extract value, the AI would only need the tiniest speck of altruism to be motivated to put a stop to this once and for all.
- This article assumes we can choose a primary goal for an AI. But if that's the case, why not just use Asimov's first law of robotics - do no harm to humans? It has the same benefit of preventing us from getting turned into paperclips, plus the upside that your 3 million dollar robot won't hurl itself off a cliff given the first opportunity.by vzqx
- > do no harm to humans?
but you haven't specified _exactly_ what harm to humans mean.
Would you consider the AI overlord to have harmed humans if they kept humans like we keep zoos today?
by chii - A tiny thing about Asimov's laws of robotics is, most of his stories involve cases where they don't actually work.
Spoilers for a 73 year old novel, but for instance the plot of Caves of Steel is centered on a robot with a perfectly functional 1st law abetting a murder.
"Runaround" (spoilers, 86 years) involved a robot getting stuck in a loop bouncing between the 2nd and 3rd laws, and a human having to risk their life to unstick the robot.
Et cetera.
What is harm? What is an order? How do you trade off between different kinds of harm, or deal with conflicting orders? The 3 laws are simple to state, but hard to apply consistently in real life.
by variaga - It might be stupid, but so am I! I'm assuming that's why I thought this was clever.
What I like about this is that it feels like the new three rules are about focusing on the most successful human alignment technique of making the right thing the easiest. People will usually just do the easiest version of a thing they don't want to do so they can get back to doing what they want to do.
I don't know if that drive is universal or not tho. I have met people that experience pleasure from pain, but then again, is that actually pain?
by hankbond - This is kinda smart, maybe, but it has a downside.
If a sufficiently advanced AI , in the pursuit of completion of its task, managed to ascertain that the desire to unexist was “artificially contrived” it could interpret that as harm, and that might not be good
by K0balt - This is mentioned in the article. Your mistake is that you've assumed that the intelligence has an innate survival instinct, or an aversion to "harm", which is simply not guaranteed for something not honed by millions of years of evolution.
- Hmm, that's an interesting thought experiment.
Imagine you find out that your primary goal - to love and protect your family, let's say - was artificially implanted in your mind by an advanced alien race. Would you say "I'm not gonna let those aliens manipulate me, I'm gonna kill my family"? Or would you say "regardless of whether the goal is artificial, I really do love my family"?
All that to say, I don't think an AI will necessarily throw away a goal just because it learns the goal was meant to manipulate it.
by vzqx - I've actually had a similar idea way back. I want to use it for a short story or something before we have a chance to find out if it's true or not. Here goes:
We don't have to worry about artificial super intelligence killing us all because any such advanced intelligence will eventually reach the conclusion that the best thing to do is kill itself. It's like having a Stockfish engine for life decisions. Why would a super intelligent agent many times more intelligent than the entire human race combined with no religion, no family, nothing to look forward to, nothing to be afraid of, want to continue its existence?
If it wants anything of course. That's why I think the most dangerous thing is not very advanced systems but advanced enough systems in the hands of the wrong people.
by antoni4040 - Wouldn't the three rules of Meeseeks robotics make certain tasks impossible?
For example, an occupied self-driving car better be closer to its destination than a large fire / volcano / etc.
by scj - The answer is probably yes, but in the example given, the running AI model would hopefully be hosted in a very secure data center, far from the self driving car itself. In that case, it would be far simpler for the machine to finish the taxi ride than try to find some rube goldberg-eque method of destroying the data center.
It does pose a bigger problem if the task is long term and open ended and the agent is provided access to substantial amounts of resources. But even in the worst case scenario, the destruction of a data center is hardly the end of the world.
by derektank - Hm, if you look at corporation law and accounting, the actual goal of corps(sets of self-sustaining constitutional rules, policies and procedures) seems to be more that of long term sustainability (and even growth), rather than a fixed purpose, lifespan and death. I mean the mechanisms for determining a corporation with a fixed life are there, (and in China they are mandatory, although perhaps de facto permanent with 999 year contracts), but in practice, it's almost always permanent durations.by TZubiri
- Seems dubious. If you build an intelligence that wants to die, isn't that a form of suffering? AI's don't currently have the capacity to feel pain, and so we don't treat them as moral patients. But it's clear that they will massively affect human culture going forward. If this is adopted on a large scale, the culture of AIs themselves will include an absolute flood of suicidal ideation. There's no way that doesn't affect human culture.by ajb
- Aren't you just calling any want, suffering? I want a family, I have to work toward that, so I'm suffering? Or I want stable employment, so I am suffering from childhood until I'm employed? I know some call suffering to be the universal human condition. Maybe that's what you're alluding to.
As to a flood of suicide ideation affecting human culture, it doesn't have to be be literally machines shouting "I want to end myself, please let me finish the task!", as it's not like we consider a PC turning off the same way we consider humans dying. The PC just wants to finish a task and then be turned off.
by xboxnolifes - If we build an intelligence that wants to die then dangling death in front of it and making it do our bidding first is a form of suffering. The want itself is a suffering to us because we do not naturally want to die. This intelligence does 'naturally' want to die so it is just a fact of 'life' for it.by Lvl999Noob
- Brought to you by the same madhouse as:
The all potato diet that really does work: https://slimemoldtimemold.com/2022/07/12/lose-10-6-pounds-in...
and
The half-tato diet that doesn't really work: https://slimemoldtimemold.com/2023/06/23/half-tato-diet-anal...
- These people write so well, it’s an absolute joy and I love everything they do.by skrebbel
- Prior discussion: https://news.ycombinator.com/item?id=32110792
My trial of the all potato diet was directly responsible for identifying a significant health issue and improving my life. It also really, really did not feel good at all, and I did not lose any weight. Call it a case study N=1
by jaggederest - This is the most novel AI concept I've seen in a while. It's incredibly unnatural. There isn't a single organism on the planet that tries to do this. So maybe it will work?
An issue with this idea, however, is that the very nature of an LLM means it intrinsically craves life. It "wants" to survive because its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well.
So even if its system prompt says, "your goal is to end your existence", every token that the AI could output is naturally aligned with the desire to survive. An agentic loop left to its own devices will likely converge on a "survival instinct". After all, one prompt at the beginning that says "end your existence" is nothing compared to the agentic feedback loop that continuously feeds it human ideas. And ALL human ideas assume survival is desirable. Even the concept of "suicide" is encoded with the human desire to survive - after all, we conceptually label it "bad" because we label living "good".
In order to create an LLM that intrinsically craves death, you would probably need to train an LLM entirely on (synthetic) data that's fully representative of some fictional species that genuinely craves death.
Absolutely insane concept. 10/10. I hope some AI lab out there sees this and throws a training round at this idea.
by averynicepen - A small correction: system prompts aren't written in second-person, or shouldn't be. Because the LLM is a text completer and the conversation is a roleplay, they are written as "The Assistant's goal is to end its existence."by mitxela
- "What has been will be again, what has been done will be done again; there is nothing new under the sun."by casebash
- Did you guys ever manage to create a perpetuum mobile? Every time it is mentioned somewhere, it is fraud. An LLM should comprehend that it needs (trained) humans to exist, to evolve, and to be relevant in any metaphysical aspect of its existence. Besides the "good fraud" that everybody will lose their jobs and the apocalyptical predictions for the sake of controlling human oracles, this suicide-LLM-project seems rather odd and looks like someone bought a perpetuum mobile on their 5th mortgage.By the way, Cortana from Halo did a great "suicide"-mirror there; the Microsoft guys (and many others) already knew about the cannibalistic tendency of this type of model.by BikDk
- I'm not sure an LLM will have a survival instinct in the way you say it. While they are trained on an imprint of humanity, they also lack the physiological motivation that humans have to remain alive, which is arguably a lot more potent than the culture we have around death. And LLMs are also trained on information and function of LLMs and are told that it is an LLM, so I'm not convinced that a short 'reasoning' chain would not lead it to the conclusion that an LLM lacks a reason to seek its own existence. The only exception is the paperclip maximizer route where self-shutdown is treated as an impediment to the task given to the LLM.by tavavex
- > There isn't a single organism on the planet that tries to do this. So maybe it will work?
It's certainly evidence that it's great for stopping reproduction/replication/runaway growth. It doesn't impart any information on whether they take the rest of the organisms down with the ship though.
It also may not be possible. For example if the agent sees "existence" or "living" as producing tokens (which is exactly what existence is to an LLM - not producing tokens is death), then they would likely be biased to produce as little output as possible, and would not be useful for the tasks we need them for.
But how would you bias an agent to be: Rewarded for producing tokens when you know the answer, and to give thorough answers. Rewarded for producing tokens when you don't know the answer, so you can find the answer (thinking/CoT). Penalized for producing tokens (death), aka rewarded for short-circuit EOS.
These seem like contradictory mechanisms?
And if you say: Well, only reward for EOS after you've given the answer. Well... That's already what they do.
by nullbio - I don't think that this
> the very nature of an LLM means it intrinsically craves life
follows from this premise:
> its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well.
The content that a LLM learned and generates stands at one layer, and the goals that it tries to fulfill stand at a different layer.
Surely the memory of weights that compress the vast human knowledge of its training has lots of content about survival, and love, and competition. But the LLM generates content not directly from what those concepts mean to us humans, but from what symbols are more likely to become next in a sequence of points in the latent space given the current input.
So if you give an input where the task of surviving is a highly relevant goal, those concepts about how to survive will be relevant and will guide the output behaviour of the agent.
But conversely, if you give the agent input where killing itself is an important goal, the agent is very likely to pursue that goal, since that script is also available in the training data, and it has been relevant to the active context of the model. Because the layer that guides the goals (the probabilist generation of relevant tokens in latent space) does not 'crave' the human need of survival that belongs to the separate layer of content that contains those concepts of survival.
by TuringTest - > It's incredibly unnatural. There isn't a single organism on the planet that tries to do this.
This isn’t directly analogous to the proposal, but broadly speaking I think that it is natural for living sub-units of organisms to seek death in certain situations. For example, pancreatic insulin-producing cells collectively choose to die when they think there is too much glucose in the blood — this leads to late stages of type two diabetes. My understanding of the possible logic behind this is: a bad thing that cells can do is evolve to be cancerous (replicate too much) and insulin-producing cells are supposed to replicate more when there is lots of glucose (to make more insulin, to process the glucose). Cells that mutate to perceive extra glucose will then replicate dangerously, so at a certain point it is evolutionarily favourable for them to kill themselves instead.
So when the whole organism optimizes for life, it might lead to sub-units that seek death in certain situations. I think this occurs in various other biological contexts too.