Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > Why we restrict model training

    > When Outputs are used to train new models without our oversight, additional risks emerge. Safety controls may be lost - ...

    ---

    > What you can do with Outputs

    > You can use Claude's Outputs to train models that don't compete with Anthropic's own models.

    Why even include that bullshit at top? It is brazenly obvious to anyone with a brain that Anthropic and other frontier AI Labs have no real way to monetize or recoup their investment unless they strongly guard the usage of their models.

  • This means you can train smaller models than Claude’s but not competing.

    I’m not sure how an AI company feels that’s safer for their business. Specialization will always best generalization trying to do the same.

    by j45
  • I had the same thought.

    "We need to control the model to protect you from evil robots, except it's fine if the evil robots are not competing with our business" is hilariously hypocritical.

  • Any AI safety experts here? I'm wondering if this claim here really holds: > Anthropic invests significantly in making Claude safe, helpful, and harmless. We conduct rigorous pre-release testing, implement multiple safety layers, and continuously monitor our models' behavior. When Outputs are used to train new models without our oversight, additional risks emerge. Safety controls may be lost – models trained on Claude's Outputs won't have our safety measures, potentially leading to harmful or dangerous AI systems.

    From my understanding distillation pretty much copies behaviour. If someone intends to distill from a Model, it can't extract unsafe behaviour, but would learn the same safety mechanisms, no?

  • Nobody owns Claude's output (nor that of any AI), they're in the public domain, they cannot be copyrighted.

    The author must be human, at least for now, in the US this precedent holds as far as I can tell:

    https://law.justia.com/cases/federal/appellate-courts/cadc/2...

  • From my interpretation, in order to get the data you 'own' (it's not theirs to give away since they can't claim the copyright on it), you need to use their services. The agreement the user has with Antropic is for the service, not a restriction on how the data is used.
  • so if i publish that data, and someone else trains on it, am i liable? am i responsible to ensure that noone trains on that? how am i supposed to enforce that?
  • not being able to train in it is a restriction on using the data though.

    if i publish that data, and someone else trains on it, am i liable? am i responsible to ensure that noone trains on that? how am i supposed to enforce that?

  • Wouldn't such language and reasoning from Anthropic be an argument that they needed written permission to train their model on data from websites?

    Has any individual somewhere around the world tested this in court by now? Sued Anthropic for copyright infringement because Claude can reproduce information that is only available on their website?

    It shouldn't be that expensive, right? If you sue them for - say - $10000 then what would the costs of such a court case be?

    Personally, I think "learning" is not a copyright violation. But if they themselves make it one, then they should also face the consequences, no?

  • Apart from the usual hypocrisy,

    > Safety controls may be lost – models trained on Claude's Outputs won't have our safety measures, potentially leading to harmful or dangerous AI systems. We also have no visibility into deployment, meaning we cannot monitor how these distilled models are used or prevent misuse.

    So was that why your models caused three real-world security incidents?

  • > So was that why your models caused three real-world security incidents?

    No, that was just marketing :)

  • Isn't that someone else's problem anyway? Someone else's model trained on your outputs is still someone else's model.

    > we cannot monitor how these distilled models are used or prevent misuse

    which is fine since they are none of your (their) business..

    I think the only possible concern is using their models name - making it clear it's a new model merely trained in another one should fix that.

  • So can you write GPL content using Claude? MIT? Because if this condition applies to the output, I don't see how it's compatible with FOSS.
  • Yes, obviously, you own the output. Training on those outputs is a violation of the service agreement, and that’s separate from who owns the outputs.
  • MIT doesn't bar additional restrictions, so you could say this is MIT except you can't train on it. GPL on the other hand does not allow additional restrictions and therefore the code would not be compatible.

    but you can't copyright the output anyways. which either makes the restriction on training void or, it means the owners of the model own the copyright, and they only transfer some of the ownership to you. is there such an ownership transfer statement? i haven't seen one yet.

  • It might depend on your jurisdiction, but LLM output is usually not copyrightable because it is not made by a human. There might be a few exceptions, if you guided the LLM in very specific and particular ways. But in general, the output of an LLM is public domain the moment it was produced.

    Can you never tell anybody and just slap a license on it? Sure.

    The issue here is not the license, but that you violate their ToS (if that is valid and enforceable is a different question of course). But if you publish the LLM output on github, and someone else takes it to train their LLM, and you did not actively encourage or help them, it's fine.

  • Corollary: If you can't use the output to train, then you don't own them.
  • Apparently, they use "own" in the same misleading way that digital media providers use "buy", but later remove the purchased content from my library.

    It's named "yours" at the time of payment, but actually only a license to use under certain conditions.

  • Yup, either you own it, or you don’t own it and ownership is being redefined as a quasi license.
    by j45
  • I was not asked for permission when they trained their model on my output.
  • They stole all the data, and then dont want you to steal it back. Its basically Robin Hood all over again.
  • Piracy is not theft.
  • I have released some of my projects as Open Source, I also have a company with privative software.

    Claude and AI partners have taken all what they could from the Open Source projects without giving credit or respecting the licenses. They have increased the traffic on websites in an absolute disrespectful way increasing the hosting cost in inefficient and ridiculous ways.

    They have taken all the important books and not asked permission from the authors.

    Fair enough. Fair use.

    Of course I would create a competitor software to Claude or any others if I could. Using Claude(and others) of course.

    I am not paying you 200 dollars/month for you to tell me that I could not create code that competes with you. If you try to go to court in Europe with this you will lose.

    It is just the same fair use you proclaim for taking the data from others.

  • They would likely lose in the US too. But until that happens, they can keep living in their fantasy world...
  • > When customers use Claude to generate Outputs that then train competing models, they're essentially using our infrastructure and investment to build direct competitors to our service

    We did so, please do not repeat it at home.

  • Careful: pointing out this distillation hypocrisy (“rules for thee but not for me”) is liable to draw moderation deeming it “a thought-terminating cliché” that is against the HN Guidelines. See, e.g., https://news.ycombinator.com/item?id=49007792.
  • Didn't they, ehm, use all the historical infrastructure for free to train the thing?