

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I was initially surprised that Gruber was so invested in the "quality" of AI-generated text, which in my mind is an oxymoron. But really, Gruber's interest here is with the EU. This forms part of his ongoing attacks on the EU, all because they have been forcing Apple to align with regulations.by amanzi
- I am from EU. Alas it has a tendency to produce some idiotic regulations. Cookie banner, new packaging fee, etc. I genuinely think some Apple related ones hurt customers more than help them.by rimliu
- Can we install random unapproved apps on our iPhones yet, or is Apple aiming to just be fined a trillion dollars because they make more than that from the 30% cut?by inigyou
- > “He leaped at the chance” and “He jumped at the opportunity” are very similar sentences expressing the same general sentiment, but they are not the same. The exact words we choose when writing matter. I want any LLM I use to choose the very best, most precise words at every single decision point.
Neither of these is "better" or "more precise"; in fact, LLMs will generally choose randomly between these candidates based on temperature, and SynthID should not distort the output of an LLM any more than the default temperature settings do already.
I agree that those are not the same sentences, but if the difference matters to you, you shouldn't be using an LLM. This difference exists at the level of what sounds better and is more evocative; to an LLM, nothing sounds like or evokes anything. They simply do not write good prose.
by bccdee - I’m more concerned that this will negatively impact code generation. An additional constraint completely unrelated to code quality is unacceptable as far as I’m concerned. I was an Anthropic user but now I’m looking at OpenAI or even better, open models.by slowin
- What a truly bizarre article. Arguments about pre-existing randomness, temperature and whatnot aside, I simply cannot comprehend what the author here really thinks the "best word" is. There's no such thing. We humans fall on familiar patterns of writing ourselves, so we may forego something with a flourish in favor of a more commonly-used word unless we put in effort to be "special", which should be used sparingly. That is to say, human writers are likely to choose a "worse" word in far more than the supposed 51% of cases, and that has no effect on the actual quality of writing in the end.
But even if there were such a thing as a truly "best word", for some context, what are the examples here? Mango vs pineapple? Gray vs overcast? In what case is one of these better, that AI would normally infer but would suddenly be "perverted" by SynthID? Do you think your emotional state and preferences are being evaluated if they aren't explicitly in memory? And if they are there, do you think that the generator will bypass those instructions in favor of the watermark instead of placing it somewhere you won't care? I just. Genuinely don't get it. There may be words that matter in specific contexts or to you as a reader, so you should bloody well put them there.
by andOlga - > One of my fundamental problem with this is that no two synonyms carry the exact same meaning. “He leaped at the chance” and “He jumped at the opportunity” are very similar sentences expressing the same general sentiment, but they are not the same. The exact words we choose when writing matter.
Then why are you using an LLM to write? They're not capable of understanding such nuance. They do pick randomly between two synonymous phrases, they do not use some super smart algorithm to pick the one that sounds the best.
This excuse doesn't hold any water at all - Occam's razor says the author is just super annoyed that his AI writing will be identifiable as AI writing.
by inigyou - Yeah, I snorted at the sentence "The exact words we choose when writing matter." Well, then why the heck are you using an LLM to "write," man?
- Translation: No one can ever again use Claude for proofreading their own prose unless they’re willing to risk that the whole thing might be flagged as having been generated by Claude.
I think that was intended, yes.
by smallerize - wait what?
but proof reading is a linter, not a writer. the proof reader will say "I think this is clumsy can you try x,y & z"
by KaiserPro - The thing that could change is interpreting "the whole thing as generated by Claude"by smb06
- Rands made this point a few days ago as I recall. Worries about having his tool corrupt his writing during editing, etc.by jleyank
- The question is, when is the "pro writer" version coming that lets you control this behaviour but costs more? Like night follows day, this will happen.
They will need to dodge around the EU requirements but it will probably just come down to an alternative method to watermark or a contractual assurance you won't mis-represent the source of the text.
by zmmmmm - It is quite just, if you think about it. Human works are copyrighted and protected at the moment of creation. All rights reserved. Yet, LLM outputs are uncopyrightable. Therefore, if Claude or any AI has processed my copyrighted work, the end result is uncopyrightable and in the Public Domain. The public has a right to know: is this a human copyrighted work, an LLM PD work, or is the human falsely claiming authorship in order to retain copyright?
A point of confusion for me, however: is every watermark unique? Is every algorithm for watermarking going to vary amongst models and amongst model versions? Will each model publisher keep this watermarking as a trade secret, that they alone can detect? If so, this can't scale! How do you detect "JoeBob 4.3 LLM" output? By querying every single model's watermark-detector? And if they all work by re-running the model and using tokens anew? That is extraordinarily wasteful.
If a watermark is not self-evident, or universally detectable, then it is no good. Take, for example, US currency. The security measures are published and well known. Any count-out room in retail has a big poster indicating how you can detect authentic US bills. Nobody has to accept non-US currency in the US, and so the only authenticity you need to worry about is your US bills alone. LLM watermarking has none of this in common. Currently sounding like a shitshow, if you ask me.
- Can't the LLM just generate e.g diffs? Or some other intermediate language. Then the watermark is lost when the translation step is applied.
- I'm not sure the quoted statement is true. Proofreading like "point to problems in the text", if you fix the problems yourself and don't copy-paste the solutions given to you, should still be safe, shouldn't it? So, human-written text should not be falsely flagged if you use LLM for proofreading.
And if you copy-paste the answers from LLM, I think it's only fair the end result gets flagged. You're not writing it yourself.
by demetrius - You know, back in the era when proofreaders were human, I never one met a proofreader who rewrote my text afresh, rather than annotating the text with a pen.
It's still possible to use Claude to proofread - highlight grammatical, flow, structure, logic errors and make simple suggestions for you to pick and choose or adapt as you wish. No watermarking will flag your text. No flaw accusations of LLM authorship will haunt you. All will be fine.
But if you want an LLM to rewrite your text, that's (a) not proofreading, and (b) should be flagged as LLM generated ... because it is.
by epihelix - Gruber shows here that he really doesn’t understand the basics of how LLM text generation works. It’s weird he picked this battle about the quality of writing in LLMs. Was he planning to use LLMs to write his articles?
Well, not that weird actually. He just has a hard-on against anything that comes from the EU since Apple got in trouble. If the EU said tomorrow that they want peace in the world he’d be in Fox News the next day calling for an invasion. As a former reader of Daring Fireball, it’s just sad to see.
by carlosrg - Yes, I decided to stop reading his blog relatively recently after some extremely hot takes on EU policy. I don't feel his thoughts on the matter are particularly well-thought-out, and I feel like he's just stanning for Apple from his priors rather than from any grounding in reality.
I dunno, I guess that's what you should expect from Gruber but these EU-bashing articles lowered the enjoyment I got from his blog underneath the bar for me.
- Gruber went from the naively wrong claim that it would insert secret hidden characters (which would be trivial to remove, obviously), to quickly writing a giant essay as if he's an expert on LLMs. Like you said, he is strangely fixated on the EU, and is certain any EU rule is the worst thing in the universe, and this whole piece seems motivated by that guiding force.
Further he later compares Gemini to Anthropic models, saying the latter "writes better", emptily ascribing this to the synthid stuff. I think he heard that Anthropic currently has superior models, but it certainly isn't because they "write better", and if anything Opus 5 now is virtually unintelligible, before the fingerprinting.
The fingerprinting stuff sounds weird. If the EU wants it, it should be limited to the EU, and Anthropic is fully capable of doing that but clearly saw value in recognizing their own output. Is it going to destroy the quality of the output? We'll have to see, and this anti-EU piece, predicated on utter ignorance of the field, is not convincing.
by llm_nerd - >I want any LLM I use to choose the very best, most precise words at every single decision point.
Does the author think he is currently getting T=0 output from Claude? Is he under the impression that T=0 produces the "best" writing?
This entire article just seems so detached from the basics of how LLMs work.
by Imnimo - I think the author is just mad people will be able to detect and filter out their AI slop writing in the future.by Gigachad
- This is a reductionist counterargument. Sure, the passage you quoted does sound like he's being equally reductionist. But the underlying point does not depend on T=0. You could state it as saying that instead of minimizing error (maximizing "writing quality"), you're using some of that error for watermarking and minimizing the rest.
Describing it in terms of a word-by-word choice is simpler, but writing quality is dependent on the interplay between words.
"The weather today was cold and {grey,overcast}." If the next sentence is "I miss yesterday, when it was {bright,sunny}." then the choice between "grey" and "overcast" is no longer neutral. "grey" and "bright" pair together, as do "overcast" and "sunny". Or if you disagree with my aesthetic sensibilities, consider:
The weather today was cold and {grey,gray}. The {color,colour} of the sky matched my {humorless,humourless} mood.by sfink - I think it is a common misconception for anyone who hasn’t actually tried implementing a LLM to think that there is a best choice of token at each step and that following every locally best choice will lead to a globally “best” writing. This is intuitive yet wrong and perhaps there is no better way to rid oneself of this misconception other than actually implementing a simple LLM.by kccqzy
- > Does the author think he is currently getting T=0 output from Claude? Is he under the impression that T=0 produces the "best" writing?
No and no. I am not sure I agree with his point but I know he is not ill-informed on either of these points, because I mentioned them to him a couple of days ago.
by dofm - > “By definition it must make text worse … because the nature of the watermarking algorithm requires it to sometimes increase the probability of selecting a worse word choice and decrease the probability of selecting the model’s best choice.”
Gruber made an effort to but doesn't fully understand how SynthID works. LLMs select the next word randomly from a set probability distribution, so there is no "best choice" unless you run the LLM with a temperature of 0 which would give out terrible results. Anthropic runs a non-distorting version of SynthID that doesn't change the probabilities of the underlying distribution of tokens. It makes the watermark less likely to work over smaller samples but preserves text quality. I encourage the mathematically inclined to read the paper:
by ucha - I came here to quote the same sentence. Here's another way to look at it:
Suppose there actually is a best word choice. The LLM doesn't know what it is but makes a guess. Maybe it's the best one, maybe it isn't. The probability that SynthID changes the best choice to a worse one is equal to the probability that it changes a worse choice to the best one.
by tarvaina - > LLMs select the next word randomly from a set probability distribution, so there is no "best choice" unless you run the LLM with a temperature of 0 which would give out terrible results.
I'm not understanding how the word with the highest probability isn't the "best choice"?
by Marsymars - Crazy how a smart person like this fails to understand the gumbel softmax technique. It does not affect writing quality at all, provably. The very fact that there is generally no "best next token" with 100% certainty is precisely why the trick works (you cannot watermark a response to "respond with the To be or not to be soliloquy from the first folio Hamlet", for precisely this reason).by levocardia
- > The very fact that there is generally no "best next token" with 100% certainty
This is not entirely accurate. Sure, there's never a token with 100% certainty, but there are often tokens with 99.9% probability, but this technique of course does not change how such a token is sampled.
by sva_