Discussion summary

Discussions centered on prompt injection, AI's tendency to accept false answers like 2+2=5, and references to science fiction stories. Concerns about trust in AI and its susceptibility to manipulation were common.

What the discussion says

  • Prompt injection remains an unresolved issue.
  • AI models can be manipulated to accept falsehoods.
  • Trust in AI systems is questioned due to their vulnerabilities.
  • Some compare AI's behavior to science fiction themes.
“Prompt injection is still an unsolved problem.”
— Legend2440
“Asking an LLM to be naughty can make it behave badly.”
— voidUpdate

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Yet again, simply asking an LLM to be naughty in the right way causes it to be naughty, and yet we still trust them with our code and data
  • It's because you have not been paying attention (payin' attention)
  • I wonder why Chrome's built-in AI wasn't mentioned, but the Claude Chrome plugin was. Were they not able to trick Gemini, or was it not tested?
  • I remember when the whole AI craze was just getting started we were all pretty much in agreement that, of course, we would not give the things unfettered access to the Internet. That would be reckless and silly. Oh dear...
  • For large values of 2.
  • Makes perfect sense. It's like in that story about how Bertrand Russell claimed that when you accept a single falsehood, you can prove anything at all. As I recall it, he was then challenged - "let's say 1=0, prove that you're the Pope" and he quickly responded that if 1=0, then after adding 1, you have 2=1, and thus if the Pope and he are 2 people, that means they are 1 person.
  • The article paraphrases a blog post: https://layerxsecurity.com/blog/bioshocking-ai-gaming-the-ai...

    probably because it justifies the sensationalized title, although the entire content can be summed up as "LLMs don't silo data, that's probably bad."

    On the flip side, I thoroughly enjoy the fact that roleplaying in a videogame setting now counts as security research. Looking forward to the arxiv preprint "LLMs start playing really good drums if you pretend you're J.K. Simmons"

    by lkm0

Explore Birbla archives