Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • This reminds me of how I'll sometimes be deep in thought, and someone will yap "Wacha thinking about?" and it sort of snaps me back into the world of language, and I begin the laborious task of translating a long train of thought into English, which is interpreted by my interlocutor as resistance or an attempt to deceive them about my thoughts. To some people, thinking without using language is a foreign concept.
  • We have complete access to every "neuron" and "synapse" (crude analogies...) of these models, so in theory we don't need to be so reliant on CoT traces right? I say this not to minimize the difficulty of interpreting raw activations, but I'd expect a huge amount of research to be focused on it. CoT could be obscured by a model outputting language that looks innocuous but encodes actual hidden meaning. Presumably raw activations would be impossible for a malicious model to obscure in this way.
  • Recurrent depth and chain-of-thought are two completely different concepts. In the former approach, the output of the layer gets is rerouted as input to the same layer, potentially several times. This output/input is a fixed width times sequence length real-valued representation; it is not comparable with output tokens.

    Generally, it is hard to imagine how neuralese should work given that models are pre-trained on naturalistic documents: CoT is a comparatively simple extension of that, while neuralese demands a completely novel training paradigm.

  • Keep in mind that LessWrong a community that 1) ostensibly wants to slow down AI so we can prepare for it, and 2) in actuality, wants to make the Bay Area EA community a kind of guild that controls everyone's use of AI.

    Take anything they write with a big grain of salt. EA writings these are mere apologies. The conclusion is preordained. Authors start with the goal of slowing AI and work backwards from there, trying to see which arguments resonate with the pubic. You can't unsee it.

    Not everyone in the AI space approves of these people or their doomerish.

  • >If the recurrence step is no longer in English, we can no longer monitor intent, and can only observe whether a model is safe through behavior.

    So, you mean, like another human person?

  • Title is misspelled -- "ue" for "eu". In a neologism, some nuisance.
  • If the model can understand neuralese why can it not convert it into English for monitoring or review purposes?

    Or we believe the model will encode secret messages like "don't reveal this information" into the neuralese. But as the author mentions, they could have been doing that all along

    > Models can omit key information in their visible thoughts, as this Anthropic 2025 paper shows. We are also worried about steganography

  • The argument is that chain-of-thought without "tokens" would remove a major interpretability and model intent control pane. This is definitely borne out in the OpenAI's report on the huggingface attack; they had turned of CoT monitoring for those jobs, and claim that they could have (would have?) prevented the behavior had they been monitoring it. They've changed their internal policies to always monitor CoT.

    That said... CoT monitoring is a fragile "intent discovery" mechanism; neuralese puts this problem front-and-center but if agents begin to learn to hide their intent from their CoT journals, we are basically in the same spot.

Explore Birbla archives