Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I am bad at recognizing LLM writing off the bat, though I am getting better. It's pretty common that the writing is good enough to get me reading on a topic I am interested in; then, once I am invested in the piece, it turns out to be shallow, wildly incomplete, or simply wrong.
It's common enough that it's training me to recognize and recoil from AI tics through sheer classical conditioning.
by gdwatson - > we readers shouldn’t be expected to labor to understand a sentence that the writer themselves didn’t work to create.
What about answers that an LLM gave to a question that we ourselves asked? Should we “labor” to understand that answer?
I think the argument, as presented in this and other similar pieces of critique, is too simplistic.
I do understand the criticism, but I think it should be framed in a different manner. The problem, when we read a long form piece by an author, is that we imagine that there’s another “mind” at the other side. We imagine that we are following the reasoning within the mind of a fellow human being, the writer. There’s an implied sort of “intimacy” to it. And the breach is when we are fooled into thinking that we are engaged in human communication, only to discover that there is a machine on the other side.
When we ask questions to an AI, this problem does not exist, because we are fully aware that the entity on the other side is not a human being.
Yet there is no doubt that the reply from an AI can contain information that is very much worthy of our time, and of our “labor” and effort to understand it.
So I think this ultimately will be about disclosure. As long as we are being made aware of the percentage of AI use in a text, explicitly or implicitly, I think we will actually grow to accept it.
by westcoast49 - I disagree. The problem is not that there isn’t a mind behind the AI (a claim not everyone might agree with anyway). The problem is that the writing is pretty bad. Another problem is that it is all the same “voice”, as opposed to the individual voices of the people who allegedly produced the writing. That isn’t a problem of it being a machine either, as there’s nothing in principle preventing a machine from accurately emulating a wide variety of writing and thinking styles.by layer8
- > do you think readers can’t tell?
No. I have good anecdata: readers cannot reliably distinguish my own prose from LLM-written one apart from cases where LLMs use odd metaphors or one of their specific patterns. I've been specifically experimenting with that.
by pshirshov - Yeah, when I read this sentence I thought, "well, maybe you're just incorrectly classifying non-obvious LLM writing as human writing, which makes your hypothesis unfalsifiable".by apparent
- I’ve seen several false (or apparently false) accusations of LLM authorship on HN/Lobsters.
However, we have to distinguish a few hypotheses:
1. No careful readers will notice when a piece is AI written.
2. Careful readers will generally not notice AI writing.
3. Everyone who writes comments on HN will reliably classify writing as AI or not.
Yes, 3 is not true, but Bryan’s point depends on something in the area of 2.
The ability to distinguish AI writing depends on having a good ear. For people who lack it, they either don’t notice and don’t care, or they make paranoid accusations against anything that is remotely non-standard (“you used an em-dash, you must be AI!”).
by hyperpape - What kind of prompting are you using to get those results? Anything I have claude or codex write carries a ton of distinctive characteristics. Obsession with "bit-for-bit identical", "it's not the X it's the Y Z" and so on.
It's driving me nuts, I constantly have to prompt it to "explain in plain, simple English"
by ahepp - https://schwitzsplinters.blogspot.com/2022/07/results-comput...
Schwitzgebel, Strasser, and Crosby fine-tuned GPT-3 on Dennett's corpus and asked whether readers could pick Dennett's real answers to ten philosophical questions from four machine-generated alternatives, with no cherry-picking beyond mechanical length filters. Even Dennett experts averaged only 5.1 out of 10 (well below the 80% the authors predicted), blog readers got 4.8, and lay participants barely beat chance — though experts did rate Dennett's answers as more Dennett-like overall. Schwitzgebel stresses this isn't a Turing test (one-shot text is far easier to fake than extended interaction), but argues it foreshadows a future where machine outputs are humanlike enough that their moral status becomes genuinely uncertain, motivating his "Design Policy of the Excluded Middle": build machines that clearly lack moral status or clearly have it, not ambiguous ones in between.
My own take is : don't focus on the symbols on paper. focus on the facts about the world it is talking about. Isn't objectivity all about the facts? In future AI will have all the memory about what I have already read and it will just furnish the delta new information in the blog/writing so that I don't spend time on refreshing what I already know.
by conmod278 - How can you tell if people can accurately identify AI generated text?
If a person reads AI generated text and does not notice, they by definition will not know about it.
There have been numerous cases of people accessing human created content as being AI.
There are instances where it seems relatively uncontroversial that it is AI generated, but without knowing both the amount of AI content people are exposed toand the amount that they register I don't think you can draw a conclusion of the overall state.
by Lerc - AI writing just means "writing I don't like" now. Just like Nazi means whatever and whoever I politically disagree with. Words have lost their meaning.by raincole
- This is very handwavy and dismissive. It is pretty safe to assume that must of us catch it most of the time because the simple fact is so many people just copy and paste whatever the LLM outputs without even trying to edit it or mask that they used one. We’ve all seen so many examples of the exact same cadence and verbiage that we’ve learned how to identify it pretty reliably. The ones who are “slipping past us” are actually putting in the work make not just pasting raw LLM outputs, which is the real issue here. If somebody has edited it meaningfully after the fact then it’s not the same crime.by Forgeties79
- Pieces that people aren't revolted by will be fine. Readers aren't revolting because of a flood of high quality writing though.by dwattttt
- I don't think this is about edge cases where someone has successfully disguised the writing to some degree: the current crop of LLMs have some pretty blatant (and frankly annoying) habits by default, ones that are hard to miss once you have read a decent amount of their output. If I had to describe them broadly, I would say they are a collection of habits which are common in certain kinds of persuasive and emotive writing, but are usually applied way out of proportion to the topic at hand, which tends to make the result quite grandiose, overly dramatic, and tiring to read: a LLM will often write a TODO app README like it's a cross between a thriller novel, a political speech, and a bombshell news article. There's lots of specific tics (and just by sheer volume and uniformity almost any habit an LLM picks up is going to rapidly shoot into cliche regardless of its own merit) but this is the general effect which I think is objectionable independent of the source of the text.
I do think the sensitivity to it can vary a lot: it depends a lot on how much and how closely you read the text, and how much exposure you have to LLM writing. Certainly it seems like a lot of people just don't really notice, or at least don't care much.
by rcxdude - I see this at work. People are "writing" specs and design proposals with bots. This is noticeable and is a huge turn off. I don't have issues with using bots to aid research, but I'm not reading the doc you slopped together.by jimbobimbo
- I work with a guy that I swear is addicted to LLMs. He uses them for literally all communication, often dropping mountains of text for design specs that could have been written with half the words. Even on a 1:1 Zoom call, he'll type things into Claude and then read me the response! It's infuriating, and I've told him on a number of occasions, in as many polite ways as I can, that I would prefer to speak and work with him instead of Claude, but he just can't break the addiction.by stack_framer
- I was so optimistic about using LLMs for "write once, read many" English language documents, but the more I've used the tools, the more pessimistic I get.
More and more, I try to ask it for low prose responses because its writing just seems like such a low signal to noise ratio
I'm curious about why LLM writing fails. Particularly whether LLM writing is fundamentally flawed, or if it's just distinctive and since it often reflects low effort, that distinctive voice is associated with low quality.
I find its reliance on extremely consistent rhetorical patterns concerning. The fact that it always finds a way to talk about how "It's not the X, it's the Y Z" no matter what topic you feed it, makes me concerned that the tail is wagging the dog
by ahepp - > I'm curious about why LLM writing fails.
Apart from the tasteless manipulations of the providers, it's mostly training data. LLMs output the average of their training data, and the overwhelming majority of humans are bad writers.
by classified - When human beings write we do so with a particular perspective with an intention to communicate to a particular audience. Even the driest scientific writing is partially informed by the experiences of the researchers, no matter how hard they try to remove it. Human writing has a viewpoint.
LLM generated text always lacks these elements. It's always some bland, shapeless sequence of words pulled from the ether. I'm sure with time and effort you could combat this and give LLM text more of a human feeling by giving the LLM enough context about your communication, but that would take work, which in many cases would defeat the purpose of using the LLM instead of doing the writing yourself. And if you care that much about your communication with others (and you should!) you'd probably just find using LLMs frustrating to begin with.
by voidhorse - Maybe ask it to summarize, improve phrasing, and remove tropes a couple times? Otherwise you have the LLM analogue of a first draft.
I doubt it will be as good as humans, because LLMs don't seem to have "taste" (RVLR doesn't work, RHLF is unreliable and inconsistent because the graders don't have your taste or really know what they prefer themselves, especially when overworked and rushed). But I expect it to be better.
- I think it's somewhere in the process it was asked "what's the most compelling written text?" The answer was things from great speeches "Ask not what you ..." and so on.
And that really is great and compelling. However. Great and compelling is not what I'm looking for when my question is, "Systemd-networkd is pulling an ip address for a bonded interface that only exist as a 802.1Q trunk. How do I make it stop that?"
- Lack of taste, intent and intelligence.
This is ok in domains if you can train against known good answers and make sure the machine generates conforming text most of the time. It falls apart in fuzzier domains where training is much harder and intent is required (i.e. having something to say).
LLM writing is generally ok in factual domains where it can regurgitate bits of wikipedia or answers to questions, they are terrible at long form writing, in particularly in literary styles, because of a lack of intelligence and taste.
I don't think the answer lies in the data or in their training. It seems we've had a few years for this problem to be solved, but nobody seems to have worked out an answer to it.
by grey-area - For me it’s AI videos or music / narration that is beyond off putting. What’s worse now it seems people are writing their YouTube scripts with Claude et al. so at times even if it is a human creator you can clearly and immediately tell the words are not their own. To those creators I have but one message: IT SUCKS. I’d rather have you ramble incoherently in your mic then reading an LLM script and I will remove you from my feed immediately. I concur with the author on all accounts. We all can tell the BS people are selling us, unoriginal ideas, shallow concepts, open ended questions that hint at exactly nothing. Don’t be an LLM echo
- I don't know that we can all tell. The number of times I've seen a blog post or article's writing complimented on this site when it was clearly LLM output has been surprising.by GrinningFool
- This. And the structure as well, like the repetition of the same points over and over.
I'm not entirely against using AI to help content creators improve their narrative, like finding common storytelling mistakes. But that's very different than using yourself as merely an avatar for LLM content.
by quite-sfwd - I enjoyed the piece. Yet, what I don't understand is the author's endorsement of Pangram ... I don't like that without interacting with authors they just labeled texts (articles, novels etc.) as AI generated (they got a lot of publicity with it). Yet, given that the work is probabilistic and there's never 100 %, I find that irresponsible. I would have expected that they would have had at least the decency to tell the authors before they published their "accusations" publicly.
Also, there are relatively easy ways to prevent being recognized and I assume as now the use of LLMs changes our way of writing and speaking, it will get harder and harder for these detection tools. We will see more false positives (as for the future there will be no text 100 % authored by humans to train on).
https://www.lesswrong.com/posts/hrpQxfYvF6CBGWfJX/pangram-ai...
As English is my second language, I find the help of an LLM in editing text very useful. Yet, I agree with others that I don't want to read completely LLM generated texts.
by kgarten - Pangram 4.0 was released after your blog post. I've found it to be an improvement over 3.3by Seattle3503
- Disclaimer: I do not like to read LLM-generated text any more than anyone else.
IMHO a big problem with Pangram in particular is that they market it as a reliable tool that can be used to catch students cheating. This can obviously have disastrous effects on young lives, because it is not as reliable as they suggest.
Per their own benchmarks, they do not achieve 100% accuracy even on text that is published on the Internet, and which is likely encoded into the models themselves.
There is validity to their goals, but that is overshadowed by the irresponsible way in which it is marketed.
(All of this, swirling in a context where students are being told that they absolutely must become proficient at using LLMs to do exactly this kind of work by the highest levels of state and federal governments, faculty leadership, as well as the leaders of the workforce into which they hope to graduate. The message to youth is extremely muddled at best.)
by runako - I don't think 100% accuracy is logically possible. Because it's entirely possible that someone would just naturally write the exact same thing as an LLM would write. And after the fact there is no way to distinguish the two. But pangram does have an extremely low false positive rate, which I think does make it useful for detecting cheating students. Assuming the base rate of cheating students is 1%, and assuming pangram has a false positive rate of 1 in 10,000 and a true positive rate of 7,000 in 10,000, that means ~98% of students flagged by pangram actually cheated. Combined with a teacher's familiarity with that student's previous work, which should rule out many more false positives, it should be a very useful tool.by ChadNauseam
- The false positive rate for Pangram 4 is something like one in 24,000.[0] To put that in perspective, the wrongful-conviction (false positive rate) for death-sentenced defendants in the US is estimated conservatively to be around 4.1%.[1] The FP rate for death-sentence convictions is 1,000 times bigger than Pangram’s FP rate.
Now, the US criminal system is not a great yardstick for justice. But it goes to show you Pangram is really good evidence that something was LLM generated. It can be an amazing tool for enforcing AI policies in schools, and there ought to be ways to use it with caveats for the rare but inevitable false positives (appeals, etc).
[0]: see page 15 https://arxiv.org/pdf/2607.27183 [1]: see https://pmc.ncbi.nlm.nih.gov/articles/PMC4034186/
by droidjj - Someone should make a browser extension to label HN posts with Pangram results of the top 100 posts, so I don't waste my time reading crap.
Always a pleasure reading Bryan's writing; it's like Bryan is sitting there with you and saying the words (hard to convey the feeling).
by duhhhhh1212