Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Strong dislike for papers that tell me what to do in the title, especially when even the paper admits a loose correlation of the intermediate tokens compared to solution correctness. My solutions work and they speak for themselves.
  • Peculiarly vocal, where were all these people when they started calling the machines computers, anthropomorphizing them akin to the original human (most often female) computers that used to run such calculations? And how dangerous the consequences, we've been dead reckoning for 60-70 years with the wrong terminology without course correction!

    Where were these vocal people when the "raster-oriented ink deposition machines" were being called "printers"? The meat or machine brains of future historians will melt because they can't handle ambiguity, a word gaining extra -yet similar- meaning! A word with multiple meanings, unheard of!

    Where were these vocal people when people started using software terminology like "executing", "calling", "throwing and catching errors", as if software were human -clownlike sure- but human?

    The danger!

  • There is a useful engineering consequence here beyond terminology.

    If intermediate tokens are not a faithful representation of the computation, then they are a pretty bad audit artifact too. We probably shouldn't be trying to make the model's internal narration more interpretable., but rather the computation around it more reproducible.

    Record the actual inputs, model/version/configuration, tool observations and outputs, then make the execution replayable enough that differences between runs can be isolated.

    In other words, don't ask the model to explain what it thought, and instead make the system able to show what actually happened.

  • Although I 100% agree that the core mechanism of GRPO is purely mechanical token-by-token probability generation, because RL only rewards exact final answers, the training forces the model to develop error-correction habits. This makes the output extremely like human thinking when solving a problem. It's like the order of the thinking tokens is what causes it to get that sweet, delicious reward, and this order seems like a reflection of the human thinking process.

    I created a flame graph classification of thinking-token phrases into setup, execution, decomposition, verification, error correction, surrender, and deliberation, or classified as steps in an OODA loop, which is more of a reach. It literally has a verification step and, if it finds an error, an error-correction step.

    If there is a verification sequence of tokens with an error-correction sequence of tokens during RL training, it will perform better; and if humans do these steps (did you proofread your reply to this comment? did you correct it?), they will perform better — which is why it is so easy to make the anthropomorphizing metaphor.

    Nonetheless, the paper is 100% correct that these machines are not thinking like humans.

    https://adamsohn.com/reasoning-grid/

    https://adamsohn.com/lambda-variance/

  • The anthropomorphization of LLMs should be discouraged as much as possible. It perpetuates bad practices and encourages the use of these bots for tasks they are not intended for (particularly as chatbots).

    Thinking traces should be treated as black boxes. There is no point in reading them. Only the LLMs’ conclusions are relevant. This is particularly true of Opus 5, which employs reasoning that seems highly questionable but very often reaches excellent conclusions (compared to its peers)

  • Coming from a more traditional stats/ML background, I try to view "thinking" traces as a way to explore the search space without getting caught in a local maximum.

    A better analogy for me is annealing; you can't cool metal down instantly or the result is brittle. You must cool down gradually, which allows the molecules to arrange into more durable structures. Random but controlled.

    In the same way, thinking traces are testing out all sorts of novel connections between tokens ("But wait..", "Actually,.."). Like a highly divergent branching mind map that gets pruned over time, rather than settling directly into the initial answer.

    Now consider human cognition. We're constantly diverging, daydreaming, playing "what if" scenarios and measuring up those ideas against our internal objective functions (proxy for reality) to see which ideas stick. Not too dissimilar. But it's hard to call what we do "thinking" either - it's the default mode network wandering.

  • > While a human may say “aha” to indicate exactly a sudden internal state change, this interpretation is unwarranted for models which do not have any such internal state, and which on the next forward pass will only differ from the pre-aha pass by the inclusion of that single token in their context. Interpreting the “aha” moment as meaningful exemplifies the long-neglected assumption about long CoT models – the false idea that derivational traces are semantically meaningful, either in resemblance to algorithm traces or to human reasoning.

    This paper addresses something that has always bothered me about LLMs. You read their reasoning, see something like “Wait, that’s wrong” and then watch them make the exact mistake they just identified.

  • Is anthropomorphizing a real problem? From what I know, none of the serious LLM researchers believe it has anything to do with human reasoning, apart from Anthropic with their click-baity terminology like "LLM biology". It's just a metaphor. "Reasoning tokens" is simpler to say than "learned prompt augmentation tokens". I used to (and still do) anthropomorphize things long before LLMs, and I've seen my colleagues do it too. Say, when MySQL fails to start because it tries to read its config from the wrong dir, I may say "oh, this guy thinks he must read the config from ..." (having a language with grammatical genders as my native language also helps make it sound pretty natural). It's more fun like that :) Doesn't mean I genuinely believe a MySQL instance actually thinks.

Explore Birbla archives