

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I thought this was about all the illegible jargonby bloppe
- Oh look! Peirceian firstness for LLMs!
- James Mickens! I was hoping for more jokes …by chubot
- I had not heard of this comic masterpiece “Pfau et al. showed that a model whose chain of thought is just dots (“...”) can nonetheless … solve problems that are intractable for a model with an equivalent architecture but no chain of thought.”by applicative
- I think the point of the article/paper is how LLMs could be saying something but thinking something different or more than they are saying. Like Anthropic's article and video about Claude's "j-space". I do agree this is a field that demands investigation because it goes beyond thinking: "ok this models should never speak in a language we don't understand.". It's fair to think they might have hidden thoughts even speaking a language we do understand.
And well if I missed the point of the article, sorry. Anyways AI should be kept understandable and as see-through as possible if it's gonna be more powerful than a human.
by bcorigliano - fundamentally, "linguistic illegibility" is a new term for something that we've known about for about a decade now. In RL the more general ideas is "reward hacking" and in NLP it has been called "semantic drift".
I dislike this term because it doesn't explain where this "illegibility" is coming from. Models are post-trained towards non-linguistic goals with (mostly) non-linguistic rewards. A model's reasoning chain is reinforced if it leads to a correct answer or agentic goal. It doesn't need to be linguistically accurate and meanings can drift over training.
by mnkv