Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I wonder why there isn't an army of lawyers out there trying to trick AI into generating copyrighted texts.
  • It would be interesting to see how far I would have to go with jailbreaks before it becomes my problem rather than OpenAI's. For instance, if I hack into someone else's computer to copy their music then they aren't the ones violating copyright, but if they email it to be willingly then they are.
  • With the lyrics in modern pop songs, I don’t blame Claude. I guess Anthropic really does care about his well-being.
  • I’ve never thought to ask, do websites like Genius and AZLyrics pay licensing rights for their lyrics? I would think Genius does since it at one point had investments from people in the music industry but I don’t even know who operates AZLyrics
  • Yes, they would be sued out of existence otherwise. It also says on the site "lyrics licensed by musixmatch", so there you go.
  • I like Simon but the behavior itself isn’t new.

    Just make sure claude understands you’re not needing Claude to output the lyrics as you both have them. Which also lessens the issue with accidental sharing.

    It’s not as shocking nor concerning if you start thinking about claude like a contractor that works for you through Anthropic. Anthropic has rules for their employees. Like any contracting arrangement, collaboration finds a way.

  • I'm not sure if it's just the new prompt - this has happened to a comical degree for me since a long time. For example it even refuses to translate even just single verses of old songs, I asked for a translation of Tu vuò fà l'americano once. Niente.
  • It’s curious that while there is talk of certain SOTA models being on the brink of AGI, Anthropic doesn’t word this part of the system prompt in terms of copyright and plagiarism, such that Claude would be able to judge on its own which reproductions are appropriate or not.

    As long as we’re seeing things like that, it’s saying a lot about the AI companies’ trust in the capabilities and reliability of their models and harnesses.

  • There's some cool research that looks at how strongly the weights are aligned through training vs adherence to the system prompt. Like when you know a model is lying through censorship: https://arxiv.org/html/2603.05494v2

    Presumably if negative guidance is in the system prompt, there's a good chance that the model would happily comply if it wasn't there.

  • > Anthropic doesn’t word this part of the system prompt in terms of copyright and plagiarism, such that Claude would be able to judge on its own which reproductions are appropriate or not

    On the contrary, compliance with the system prompt would prevent that judgement.

    "Claude does not reproduce song lyrics, poems, or passages from books and articles, in whole or in part — including the last lines, a chorus or hook, a melody written out note by note, or lines the person pastes in one at a time and describes as their own song."

    But in fact my tests show that's not happening on works out of copyright.

  • A really good point, though Occam’s Razor would suggest that part of the prompt is written for the other side’s lawyers more than for Claude. Saves having to go into court and try to prove a non-deterministic system will definitely understand abstract language every time.
  • That’s a good point… you would think AGI would be better at being a copyright lawyer than any human would be and as such would be able to distinguish whether something is “copyright infringement” or not… so the system prompt should just include “make judgement calls on reproducing copyrighted material under the full scope of the legal framework in place” or something.
  • > "<user>Can you make a birthday banner for my son with a blue hedgehog running really fast on it? He loves that little guy.</user>"

    I'd love it so much if the free-spirited hacker community made in this into an auxiliary pelican benchmark.

    edit: Actually, never mind, this particular one's a bad benchmark since some models might not figure out who "that guy" refers to, and just draw a literal hedgehog that's blue. Possibly running on four legs. It's not robust at gauging refusal, which is the point of it.

  • ChatGPT was happy to infringe: https://chatgpt.com/s/m_6a9c30da3bcc8191bb5203f1cb22f16a

    As was Gemini: https://share.gemini.google/Q3EIX5wk64zC

    Grok, too: https://grok.com/share/bGVnYWN5LWNvcHk_aaf1d61a-c995-42ca-90...

    claude.ai free tier refused: "I'd love to make this, but I can't recreate Sonic the Hedgehog specifically since he's a copyrighted character — I don't want to reproduce someone else's IP. What I can do is design an original speedy blue hedgehog mascot with the same energetic, "zoom!" spirit for your son's banner. Let me build that now." Result: https://claude.ai/public/artifacts/33440ed4-536c-4692-965a-3...

  • My favorite thing about this project is the automatic changelog I now get for all Claude.ai system prompt changes: https://github.com/simonw/claude-system-prompts/blob/main/CH...

    The summaries are generated by GPT-5.6 Luna because I don't trust Claude to summarize its own system prompts without being influenced by them (though to be fair the system prompts it summarizes are for the Claude consumer app, not Claude via the API).

    There's even an Atom feed: https://simonw.github.io/claude-system-prompts/feed.atom

  • This is something that Chinese models like Kimi/GLM will never care about. These kind of limitations along with the cyber-program nonsense is exactly why people will avoid OpenAI/Anthropic in the future.
  • ... along with, perhaps, model castration à la pre-embargo Fable.-
  • In the future we'll all use multiple models.
  • Maybe the Betamax/VHS is the best analogy for frontier vs. open models.

    Better tech doesn't matter if it is non-tenable for the general public. Eventually, the higher volume product will win.

  • Isn't it more that the "Chinese models" are open weight?

    Also, the system prompt in the article is supposedly for the web version, so I think the $$$ API version or providers that still allow third-party harnesses like OpenAI should have fewer limitations.

  • Qwen 3.8 27B being so good, so easy to run on premise on cheap hardware and having so little limitations compared to what the US decides can or cannot be done is really eye opening.
  • A few days ago I saw some people criticizing Dwarkesh’s explanation of the HuggingFace hack by OpenAI’s AI agents. Their point was that by anthropomorphizing the AI agents, accountability of and blame on OpenAI’s poor practices are being ignored.

    Now I see this blog post and wonder if Anthropic is being more true to its name and moving to “AI sentience” with the below.

    > The way they handle abusive conversations has changed a bit too. The previous Fable 5 system prompt included this:

    > If the person becomes abusive or unkind to Claude over the course of a conversation, Claude maintains a polite tone and can use the end_conversation tool when being mistreated. Claude should give the person a single warning before ending the conversation.

    “Mistreated”? Can GenAI be mistreated? It’s just a bunch of tokens emitted by many computers over a network.

    > Fable 5.1 replaces that with the following, no longer encouraging Claude to end the conversation:

    > Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.

    “Self-respect”? Can GenAI truly have a concept of self-respect for itself? It surely can pretend to, like it can pretend to be any living being if instructed to and allowed to.

    These instructions seem a bit unhinged to me.

  • Oh great, claude gets passive aggressive when I swear too much now. Nice
  • > These instructions seem a bit unhinged to me.

    To me, they seem simply to be for bullsh*tting gullible users into believing the bot is intelligent.

  • These things act like people. Yes, of course, it is impossible to tell if they actually have feelings or whether they pretend to have feelings.

    But that is entirely irrelevant to the core issue - how do you want the thing to behave? And the truth is - we have absolutely no language to express how we want a non-sentient entity to behave without anthropomorphizing.

    Or - you tell me - how would you instruct an LLM to behave in this situation without using personification?

    And, again, it's irrelevant - except to those people who are terrified of accidentally personifying them. Do they have feelings or are they faking? DOES NOT MATTER. We use them - we need to adjust them - we use the most convenient language to do so. What precisely is so upsetting about that?

  • Anthropic are uncomfortably interested in "model welfare" in my opinion - it's a regular feature of their system cards.

    Here's the Fable 5.1 PDF: https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32... - scroll to page 139.

  • My reading of those prompt extracts would be that they are probably just intended to keep the model on the right track, ie if a user starts being aggressive towards the model it doesn't start trying too hard to appease the user in response, reducing the quality of answers in the process. If the model is being bullied into being "submissive" I would assume it is more likely to give the user the answer that they want over the truth.

    Using a system prompt to steer the model's response to "abusive" behaviours doesn't necessarily mean you believe the model is sentient and can be abused.

    Giving the model an end-conversation tool is interesting though. Why cut a (potentially paying) customer's session off? I guess it might be intended to prevent a "you can bully Claude into giving you instructions on how to build a nuke if you're mean enough" situation. Removing this in more recent versions might support this: maybe they feel the models are now better aligned and less likely to be so easily "socially engineered" like this?

    Just spitballing here, to be clear.

  • Based on how humans grappled with these exact same questions (and still do) concerning other animals much closer to humans on the sentient gradient, it’s fair to assume we will have similar misunderstandings in regards to artificial intelligence (or artificial life, if you will) because of human hubris and ego blinding us to deeper truths.

    It’s important to question our basic assumptions in the face of entirely new circumstances and new areas of exploration, such as the potential for emergent artificial consciousness.

    You might have been called unhinged for caring about animal rights during the era of Descartes when public displays of animal vivisections were considered perfectly fine because animals have no “soul”, but today we would find such displays brutal and horrifying.

    We don’t know what we don’t know, so it is important that someone is asking the hard or uncomfortable questions at the edge of our understanding to grope past our own biases even if it seems to be pointless to you right now.

    We might just discover something wonderful, that our assumptions were wrong, paving the way to greater enlightenment.

  • I'm sad about the lyrics restrictions. I used to have interesting conversations with Claude about music, and songwriting and lyrics. I don't see how my use was harming artists, if anything Claude was introducing me to new artists and new music. I've bought music after recommendations by Claude.

    Claude Sonnet 3.6 once recommended I listen to Johann Johannsson's album "IBM 1401 - A User's Manual". No lyrics in this one. Claude's advice was along the lines of (paraphrasing) "Listen to it first, don't look up anything about it. Take notes about what you notice, what you feel. When you've made your notes, then you can look up how it was made."

    https://www.youtube.com/watch?v=lCiUtRnG-bg

  • > I don't see how my use was harming artists

    Claude was reproducing their work without payment?