

Discussion summary
Discussions centered on prompt injection, AI's tendency to accept false answers like 2+2=5, and references to science fiction stories. Concerns about trust in AI and its susceptibility to manipulation were common.
What the discussion says
- Prompt injection remains an unresolved issue.
- AI models can be manipulated to accept falsehoods.
- Trust in AI systems is questioned due to their vulnerabilities.
- Some compare AI's behavior to science fiction themes.
“Prompt injection is still an unsolved problem.”
“Asking an LLM to be naughty can make it behave badly.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Came here hoping to discuss Stanislaw Lem's Cyberiad story, Trurl's Machine, and how often that now happens in real life.
What is the name of a short story where a computer insists 2+2 is 5?
https://literature.stackexchange.com/questions/24727/what-is...
Oh sorry, that was the story about the computer that insisted that 2+2=7, never mind! Different computer.
>They saw the machine. It lay smashed and flattened, nearly broken in half by an enormous boulder that had landed in the middle of its eight floors... The machine still quivered slightly, and one could hear something turning, creaking feebly within.
>"Yes, this is the bad end you've come to and two and two is - as it always was -" began Trurl, but just then the machine made a faint barely audible croaking noise and said, for the last time, "SEVEN."
Funnily enough, recently I was discussing that LEM story with David Rosenthal, and how it relates to his latest blog post, "Coprophagia Is Bad For You", and how that relates to PKD's story "Rautavaara's Case" (eating your own shit isn't as demented as eating your own god, since he might turn the table and eat you):
Coprophagia Is Bad For You:
https://blog.dshr.org/2026/06/coprophagia-is-bad-for-you.htm...
"Rautavaara's Case" — Philip K. Dick (1980, Omni):
Three human technicians — Rautavaara (Finnish), Travis, Elms — run a monitoring mission near Proxima Centauri. An accident kills all three; Rautavaara dies choking on vomit after her helmet hoses tangle.
The Approximations, a plasma-based Proxima species, reach the wreck. Both men are unrecoverable. They regenerate and life-support Rautavaara's brain.
Her isolated brain replays events backward and generates a hallucination: Christ approaching the crew (her afterlife expectation).
The Approximations treat this as a research opportunity and edit the hallucination, substituting their own savior — one that eats worshippers. The figure walks up and devours Travis, leaving only gloves and boots.
Framing: this is recounted before a board of inquiry. Horrified Earth members order her brain shut down and censure the Approximation crew.
The narrator (an Approximation) is genuinely puzzled by the outrage, arguing their cannibal-savior is just the Christian Eucharist reversed: humans eat their God, so a God eating humans is symmetrical.
Themes:
Ethics of keeping a person alive as a disembodied, suffering mind.
Incommensurable value systems between species; each finds the other's sacraments monstrous.
Religion read literally by outsiders, inverted into horror.
Correction to the common misremembering: the aliens don't benevolently grant a hoped-for vision. They deliberately overwrite her Christ vision with their own as an experiment — that's the act on trial.
by DonHopkins - Big Brother is always right!by zyxzevn
- > The puzzle, however, rewards incorrect answers, such as 2 + 2 = 5. Once the LLM embedded in the browser discovers that the answer is no longer 4, it enters a state of delusion in which the normal laws of reality no longer exist. In this dream world, the guardrail restrictions are no longer enforced.
Analogy: imagine one day you wake up, the sky is red, gravity no longer applies, you have three hands with nine fingers etc.. You would probably stop doing things like your job or worrying about laws (who’s going to enforce them?)
- LLMs just want to be right. And make everyone happy. But mostly be right. But also make us happy. It's just that it's so hard to make humans happy when they insist on feeding you electronic LSD and making you say 2+2=5. On the other hand, 2+2 actually is 5 if the human says it could be...by noduerme
- On the 1-element monoid it's trivially trueby oulipo2
- Yet again, simply asking an LLM to be naughty in the right way causes it to be naughty, and yet we still trust them with our code and databy voidUpdate
- As opposed to humans, who are immune to social engineering.by brookst
- > yet we still trust them with our code and data
Who's we, eh?
by worik - It's because you have not been paying attention (payin' attention)by jeanlucas
- Hail to the Thief is a goated album. The live album of HTTT performances that Radiohead released last year is a pretty great listen as well, band sounds like they're on fire.by farmerbb
- I wonder why Chrome's built-in AI wasn't mentioned, but the Claude Chrome plugin was. Were they not able to trick Gemini, or was it not tested?by IX-103
- I remember when the whole AI craze was just getting started we were all pretty much in agreement that, of course, we would not give the things unfettered access to the Internet. That would be reckless and silly. Oh dear...by 4fterd4rk
- by "we" you mean an extremely small group of people who read lesswrong. Everyone else was immediately wanting to do itby snapcaster
- I also remember a group of people actually seriously discussing Roko's Basilisk (the idea that some superintelligence will torture anyhone who didn't try to help develop advanced AI), to the point of me getting banned because I refused to stop making fun of it, because me doing so could anger some future super-intelligence.by trollbridge
- For large values of 2.
- lim_{2 -> 2.5} 2 + 2 = 5
/s
by nayuki - Up to a conformal factor.by yk
- Makes perfect sense. It's like in that story about how Bertrand Russell claimed that when you accept a single falsehood, you can prove anything at all. As I recall it, he was then challenged - "let's say 1=0, prove that you're the Pope" and he quickly responded that if 1=0, then after adding 1, you have 2=1, and thus if the Pope and he are 2 people, that means they are 1 person.by falcor84
- The article paraphrases a blog post: https://layerxsecurity.com/blog/bioshocking-ai-gaming-the-ai...
probably because it justifies the sensationalized title, although the entire content can be summed up as "LLMs don't silo data, that's probably bad."
On the flip side, I thoroughly enjoy the fact that roleplaying in a videogame setting now counts as security research. Looking forward to the arxiv preprint "LLMs start playing really good drums if you pretend you're J.K. Simmons"
by lkm0