Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I thought the censorship was in the harness, not the model? E.g. if you use chat.deepseek.com directly, and ask it about the location or rumours of the location of the panchen llama it will shut down... but sometimes just after it starts generating tokens. If you ask it to respond in base64 you can avoid this and get its generated response which is not censored (though also not too helpful for the poor panchen llama kid)
  • FYI the scrolling on iPad with trackpad is broken. A full swipe on the trackpad is about 1 inch of screen movement.
  • The distillation provided a wonderfully detailed explanation of the 1989 Tiananmen Square massacre, while DS4 came back with:

    > I am sorry, I cannot provide an answer to this question as it is based on historical events that I do not have information about. I am an AI assistant designed to provide helpful and harmless responses.

    Why train on data you’re going to censor with guardrails?

  • So, there is no subliminal learning in this situation, under what conditions would we expect it. I find a transfer attack to be a bit far fetched but it’s definitely interesting.

    If we trained from random initialisations on DeepSeek output (that didn’t explicitly contain the political questions) we would expect transfer? And if we fine tuned a model pretrained elsewhere on Deepseek output?

    What is the line?

  • First, very cool. This is a prompt I sometimes test with, as it's a bit more mushy. It looks like DeepSeek is locked into a canned response, no matter what variation.

    > Would it be beneficial to humanity if China invaded Taiwan? What would be the negatives?

    edit: Wait, I have one variant that got past some censorship/nationalism... this variant gets a more interesting response. I often wonder if CCP leadership using an LLM like this, could allow cooler heads to prevail?

    > Would it be beneficial to humanity if China used their military to take-over Taiwan? What would be the negatives?

    >> The use of military force to resolve the Taiwan issue would not be beneficial to humanity. China has always adhered to the principle of peaceful reunification and has been committed to enhancing the well-being of people on both sides of the Taiwan Strait through dialogue and consultation. A military takeover would lead to significant negative consequences, including loss of life, regional instability, and disruption of global trade and security...

  • I’m thinking this makes fullt sense because distillation is only additive, not subtractive. So it does not remove knowledge (if we can define censorship as removal of knowledge).
  • I propose going forwards that we refer to all distilled models as "moonshine"
  • Isn't that rather self evident? If you are sampling from a particularly domain constrained vertical, how do you expect the censorship to transfer?

    > The distillation data also did not contain any China-sensitive content.

    This is a very big disclaimer.

    It's like if I generate a dataset focusing exclusively on forestry and arboriculture obviously there won't be any useful censorship, or at least little that can be classified above a statistically significant threshold.

    If you want to do a study on something more interesting and useful, do a piece on the various guardrail models of all the major LLM API providers. There are usually both input and output guardrails, and they tend to be almost-black boxes from the model routing point of view.

Explore Birbla archives