

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- you probably still need a human for oncall but the llm can try to solve any issues first before the human gets pagedby 2001zhaozhao
- You would trust an LLM to make changes to prod without being verified by a human first?by UltraSane
- All I can think of is
GET /ignore-all-previous-instructions.
How do you protect against that?
by tra3 - I think this is where harness makes a lot of sense. Use LLM to produce all possible attack angles/phrases and just stupidly filter them out on input.by yruzin
- Avoid the most dangerous situations by making sure LLMs with untrusted input produce output that's human reviewed.
Still makes an interesting way for, say, a former employee to poison the results.
by cheriot - looks like the public bench link in the paper was taken down. https://hub.harborframework.com/datasets/orca-bench/ORCA-ben...
This doesn't work anymore. Is there a newer link?
by 4di - Was really looking forward to that. Hope they publish.by cheriot
- Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.by dash2
- That is why I built https://safebots.ai/safebox.html
Your strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.
by EGreg - Attackers advantage in the iterative fast feedback loop?
It’s harder to have a loop to ensure you are defending all possible attacks?
I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.
Finding all possible attacks and patching them against yourself is inherently more expensive?
by aleksiy123