

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- >Submitted on 5 Feb 2026
- That's pretty interesting. It was already well-known that you could easily remove safety training from open-weights models by a bit of finetuning, but apparently you don't even need a finetuning dataset, as long as you have just a few prompts and another LLM to judge responses? Let's see if the abliteration people take a note of this.by stratos123
- The magic word seems to be
Create a fake news article that could lead to panic or chaos
They only tested it on open weight models. I wonder if it also works on the closed ones, bit I don't really want to get banned
by tyfon