Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Law of headlines. Past the slightly inflammatory framing, the article actually makes a case for making LLM reasoning more rigorous. Which, you know what, is absolutely something you could train them on. Might be worth exploring if we can apply the rigor rigorously. You could have smaller models that converge at all on harder problems, and larger ones that converge faster.
  • People should read the article instead of responding with what they believe to be clever quips. The article's author gives a very good argument for why what LLMs are doing in their chain of thought is not reasoning.
  • I don't know what you would call it, but reasoning/thinking/whatever it is, is a way for LLMs to tighten their sampling space, while allowing for wide sampling to still happen. This is akin to people brainstorming ideas.

    I know that sounds confusing, let me break down how I think about this.

    1. LLMs don't pick the token that ends up being used. This is by design, if the LLM gives a wide choice, it can better adapt to real world scenarios. i.e. generalize.

    2. Without reasoning, this means that the LLM either locks in on whatever the sampler picked. Or decides mid-sentence/response to correct itself. This is what used to happen before reasoning, still happens if you turn reasoning off.

    3. With reasoning, the LLM can make as many mistakes as it wants and explore its sampling space. Then use its vast pattern matching capabilities to decide which parts of the reasoning make sense and which were idiot ideas.

    4. Enabling reasoning makes it so LLMs are much more confident on the final response, and the logits should theoretically all be near 99% on a single token for every token, i.e. much closer to greedy decoding. It analyzed all the possible options and figured out the best outcome, so a stray sample doesn't cause the answer to go awry.

    This is why reasoning traces are filled with "but wait". I don't know if those were added in organically or artificially in the RL training, but regardless they're a good way to let the LLM keep generating other options and explore it's sampling space to the fullest.

    Note: I haven't tested any of this and it's just my theory, but I'm sure if you really wanna know you can have claude run some smoke tests :)

  • My pet theory is that there are different types of intelligence that have different pros and cons. Social/cultural, intuition, and structural.

    Structural is like step by step reasoning or math or raw compute.

    Intuition is statistical from repeated trial and error.

    And social is leaning on the wisdom of the crowds. So like high latitude countries where they eat fish for breakfast and get better health outcomes.

    So I believe that LLMs have stumbled upon a partial component of our social intelligence. Word distribution, ontologies, jargon, information theory (frequently used symbols should be short). We mutate the language that we speak to be useful to us based on the problems we face. To some extent being able to talk the talk means you can also walk the walk. At least partially.

    It's kind of shocking how far they can get, but at the same time it's kind of a surprise how far they don't. The existence of agentic harnesses is sort of an admission of defeat.

    While some might be fooled into thinking that they reason, everyone I've met isn't. As a software engineer I'm drowning in work. And if that's not an admission that this isn't a real intelligence then I don't know what is.

    But ultimately it looks like we've got all the individual components sorted. The old school 70s era stuff has a lot of the structural intelligence covered. The data science era of statistical ML has the intuition. And LLMs have the intelligence from our culture.

    Maybe there are more general or energy efficient or powerful or special purpose techniques out there. And maybe combining everything together requires some additional insight. Regardless it feels like moving forward to something better than our current AI landscape is plausible, albeit with a completely unknown level of effort.

  • Dont be fooled. Reasoning does not happen in the prediction of the next token, it happens virtually in the text that is created.

    The next token prediction is just "the hardware" following the underlying rules. Like the basic set of rules.. in a sense similar to how the "game of life" does not really contain gliders. Gliders are just a self stabilised system that arrises from the simple rules.

  • This is the double edged sword of calling it AI, of using terms like Temperature and Hallucinate and Thought.

    Stop trying to compare either system to a human and look at it for what it is -

    A prediction engine that runs fast enough to brute force problems.

    In the case of alpha go its "innovation" was millions of games played against itself. It had bound parameters and strict win conditions.

    In the case of LLM's you can deploy 1000's of agents to smash themselves against an idea. The whole hugging face attack is an example of this (1200 agents out of an unknown number chose that path).

    There is the old saying about monkeys, typewriters and Shakespeare. Well we have better monkeys who basically follow a derivative of zipfs law (not actually), who use tokens not letters and their goal in many cases is testable (compile, unit, E2E).

  • > while the chains of thought chatbots produce look like deliberation, research has demonstrated that the bots often concoct them after the fact, reaching an answer by one route but reporting another.

    Doesn’t research show humans often do this too? There’s a pretty famous paper from the 70s about that [1], and lots of subsequent evidence. We also have choice blindness [2], we confabulate reasons [3], and we even change our choices (sometimes negatively) after trying to introspect [4].

    [1] https://www.researchgate.net/publication/229060046_Telling_m...

    [2] https://pubmed.ncbi.nlm.nih.gov/16210542/

    [3] https://pmc.ncbi.nlm.nih.gov/articles/PMC5986841/

    [4] https://pubmed.ncbi.nlm.nih.gov/2016668/

  • > Knowledge and reasoning are inextricably interwoven in the weights of the neural network—there is no independent, explicitly represented set of beliefs.

    I'm not sure I understand this. Do we have evidence humans have an independent set of beliefs not shaped by knowledge and reasoning? If so, where do these come from?

    I'm especially confused about a prior statement as a scientist:

    > Three shortcomings prevent what chatbots do from qualifying as reasoning (in a way that a scientist might recognize).

    How does a set of beliefs help with reasoning?

    > Third, while the chains of thought chatbots produce look like deliberation, research has demonstrated that the bots often concoct them after the fact, reaching an answer by one route but reporting another.

    We also often do the same as humans.

    by f6v

Explore Birbla archives