

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Related ongoing thread:
How concerned should we be about Astra's recurrent architecture? - https://news.ycombinator.com/item?id=49553321
by dang - Is there a more technical source that tells what is being done? I assume it's about reasoning in latent space e.g. the coconut paper[1]?
I disagree with the "experts" there, as if latent-space reasoning will only cause proper interpretability research rather than taking the CoT as gospel[2].
Give me that any day over the constant "you will be abolished to the permanent underclass!" talk coming from billionaires.
Edit: (A better discussion is apparently here[3])
[1] https://arxiv.org/pdf/2412.06769
[2] https://thezvi.substack.com/p/the-most-forbidden-technique
- Clickbait. Recursion in layers wont effect CoT.by dimatter
- TBH sooner or later explaining the opaque recurrence with extrinsic logic will be more sustainable, and already there is a need for that with existing models. Human language readable chains of thought are just a false sense of security, and open models already emit pretty unreadable ones at times while doing the right thing based on it (since policy optimization loops with synthetic data through reinforcement learning).
- Relying on CoT (and, similarly, asking the model to explain its actions) for analysis has always seemed pretty silly (naive?) to me, but maybe I'm missing something. Why do are we so committed to preserving it? And do non-textual models provide some similar form of tracing?by percentcer
- Anthropic seems like they, at least on some vague philosophical level, have the opposite view. They are very explicit about trying not to let the CoT enter directly into their RL process, and they encrypt the CoT traces (and recently further nerfed their API surface) to make it as difficult as practical for their customers to have any idea what their model is thinking.by amluto
- It's a matter of time until cot will happen in latent space. It just makes more sense. We as human don't do all of our thinking in words.by resiros
- I thought it was pretty well understood amongst experts that trying to interpret so-called "chain of thought" is anthopomorphization and that the actual "thought process" of models is already opaque.by ionioagnio