Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > So, the whole idea here is that we increase the effective depth from 22 to 44 block applications without adding another set of transformer weights.

    From what I gathered, LLM inference is bottlenecked on memory, right? Which implies there's "spare" compute we haven't been using? Does reusing the weights like this allow us to utilize it? (Do more math per unit of memory?)

  • What a clear and well-written article. I have only a basic understanding of LLM architecture and was able to follow along and gain intuition the whole time!
  • Probably off-topic. Astra has been kind of weird. Like, I can't trust it, weird. It has an interesting tone, especially in Codex, that is off-putting. It's over zealous at times (which is why I stopped using Claude) and gets too creative when doing agentic system level stuff. Accessing files and doing things it shouldn't do. If OpenAI was chasing Claude's approach, then they are going in the wrong direction. OpenAI has always been the "business and boring approach", which was its selling point and why I have stuck with it. Claude was always the radical one (powerful, but radical).

    Also, Astra overlooked, in my opinion, a serious flaw in its approach for something I was working on recently, which really surprised me.

    Reading between the lines, there were some breakthroughs with Astra, which I'm sure is why OpenAI released it so quickly after Sol, but probably not in the ways the traditional OpenAI customer wanted.

  • Everyone interested in LLM internals should read Sebastian. He's great.

    The tldr here is that the recent "The Information" article[0] reporting GPT 6 Astra was using “recurrent depth” or “looped transformers" made it sound like it was some special new scary thing ("secret technique!") that made train-of-thought monitoring harder to do. In fact, it's just the same as stacking more transformer layers, except that you reuse the weights and so save GPU memory. It's still just producing one token at a time, and the token sequence positions aren't interacting in any "recurrent" way that's different from a regular LLM architecture.

    So, you can still monitor train of thought with these models just fine... well, if you're OpenAI, anyway. Users haven't been able to see an unsummarized trace since o1 days, because the labs are worried about distillation of their models by Chinese labs.

    (There are some legitimate interpretability concerns about stacking transformer layers endlessly, but we're known about that for a long time. And the "looping" here isn't really the source of any new issues here, except insofar as it's a cheap way to add more layers.)

    [0] https://www.theinformation.com/articles/secret-technique-beh...

  • The MSPAINT computer use demo made my jaw drop.

    I guess it's not too different from the SVG pelicans, in terms of what it's doing, but it's still amazing to see it working in real-time like that.

  • This article is not up-to-date. There have been various benchmarks (some of which published and acknowledged by OpenAI, see the charts in this thread: https://xcancel.com/tomekkorbak/status/2095596839886274689) showing GPT-6 Astra is much less monitorable. The most recent third party benchmark I saw is showing a huge jump in capability for multi-hop reasoning without chain of thought: https://www.lesswrong.com/posts/FsCkkoGsNmPzFKRhg/gpt-6-astr...

    I don't think this is explained by the model simply being more capable and therefore achieving more per token: the usage of recurrent depth (Neuralese) is exactly predicting less CoT monitorability even at equal capability.

  • If you loop an entire transformer model on itself, that seems like by-definition hidden reasoning.

    If the output of the model is its reasoning trace, and you simply feed that back into the model again at inference time instead of outputting it - then it is by definition hidden (but I would expect you could pull both this trace and a further-down final output trace out)

  • For the research focused, there are some references in my blogpost here on what kinds of computational problems minimally require how much CoT to solve: https://blog.wtf.sg/posts/2023-02-03-the-new-xor-problem/

    Notably Will Merrill's work: https://arxiv.org/abs/2310.07923

    As for how universal transformers (looping transformers, but everyone has since forgotten prior work) will affect this, Will Merrill (again) has a paper here (https://arxiv.org/abs/2503.03961) that discusses exactly this.

    The original universal transformers is called "universal" because if you allow for per-token looping decisions, it can theoretically be Turing complete without needing CoT (some nuance here about levels of precision used).

    As for whether having little or no CoT is "unsafe": It isn't clear that the model's CoT reveal how they actually arrive at the answer. As an example, what if they provide an answer before the CoT? (https://arxiv.org/html/2603.01437v2) If this is already in question, we shouldn't be relying on the CoT for monitoring the model's reasoning.

    As always there is a lot of nuance to the topic once you get your hands dirty with the details.

Explore Birbla archives