Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • It's not only that it'll become more difficult to monitor a single LLM but that also all the instances share exactly the same "collective unconsciousness". Like you develop some paranoid gibberish fantasy language that over time only you understand - except that there are a million copies of you and all of them understand every single nuance of your gibberish.
  • A response to this, which I think is reasonable. Its a bit of a fuzzy line, but the likelihood there is "magical maniacal planning" happening here is unlikely (about as unlikely as that planning happening in hidden vector states between layers).

    https://nonlineartransform.substack.com/p/relax-about-neural...

    There's also an argument here for why its _better_ for monitoring (because we have the whole state space).

  • Scott Alexander article from 2022, "ELK And The Problem Of Truthful AI"

    https://www.astralcodexten.com/p/elk-and-the-problem-of-trut...

    It discusses "Eliciting Latent Knowledge" which is "a technical report / contest / paradigm run by the Alignment Research Center". The research investigated whether it would be theoretically possible to build a "trustworthy" AI to interpret the thoughts of another AI.

  • This reads like typical LW FUD.

    1. CoT is a pretty mediocre debugging proxy due to multiple phenomena, even if you train it for readability. You can have some direct intuition about how misleading it is by trying to hijack it to follow a fixed plan. Having a trace is useful in many cases but it's very far from reading model's intent.

    2. Nothing prevents you from probing and interpreting the state directly if you want. Passing everything through the token transport is not really required for that.

  • Lesswrong is just so comically collectively incorrect.

    This makes no sense. There is no guarantee that reasoning aligns with an outcome. Reasoning is effectively saying with more compute and ability to change attention on the fly by altering context, we can come to better answers. Reasoning often has a nice property that the English intent is aligned with what the model wants to do. But it is trivially true that you could train a model that does the opposite of what it says, or something completely random. The interpretability is incidental.

    This is why we should not particularly care if we go from one clanker blackboard to another; just choose the best thing.

  • > But it is trivially true that you could train a model that does the opposite of what it says, or something completely random. The interpretability is incidental.

    That's like saying conversational question-answering is incidental to the RLHF post-training.

  • As I said multiple times in the linked post, everybody knows that CoT monitoring is imperfect.

    The question is whether the benefits are high enough that you should be happy to drop an imperfect safety mechanism in the hopes that some day, one day, you'd get a better one. Seems like the wrong tradeoff to me but regardless, arguing that "CoT monitoring is imperfect therefore we should drop it" when we do not in fact have a better mechanism in place is silly.

    This is not how we do things in any other engineering discipline or risk-mitigation system.

    It's like saying we shouldn't have rapid antigen tests because they have nonzero false negative rates, or not writing software tests because the tests never catch all bugs.

  • We have complete access to every "neuron" and "synapse" (crude analogies...) of these models, so in theory we don't need to be so reliant on CoT traces right? I say this not to minimize the difficulty of interpreting raw activations, but I'd expect a huge amount of research to be focused on it. CoT could be obscured by a model outputting language that looks innocuous but encodes actual hidden meaning. Presumably raw activations would be impossible for a malicious model to obscure in this way.
  • but even assuming you get complete access, recovering this is, by construction, even more of an np-hard problem than the inference pass itself
  • I agree in the long run whitebox interpretability would be better than CoT monitoring but the technology is very much not ready for it today (and it's not clear we'd solve enough of interpretability before the AIs have actually scary capabilities).
  • Recurrent depth and chain-of-thought are two completely different concepts. In the former approach, the output of the layer gets is rerouted as input to the same layer, potentially several times. This output/input is a fixed width times sequence length real-valued representation; it is not comparable with output tokens.

    Generally, it is hard to imagine how neuralese should work given that models are pre-trained on naturalistic documents: CoT is a comparatively simple extension of that, while neuralese demands a completely novel training paradigm.

  • You can back-propagate through the CoT iterations or recurrent layers, same as you can back-propagate through normal intermediate layers.
  • The recurrent depth sounds a lot more like what is described in https://dnhkng.github.io/posts/rys/ - a way to add depth to a network without increasing the number of parameters.

    e: While the actual CoT in neuralese paper is Facebook's Coconut https://arxiv.org/abs/2412.06769 - not sure if any production models use that one.

  • >Generally, it is hard to imagine how neuralese should work given that models are pre-trained on naturalistic documents:

    If you look at any paper/blog etc detailing Reasoning RL runs, they'll tell you the same thing. 'Thinking' text trends towards unreadable gibberish (for humans) unless you reward for it. Even then, take a look at the scripts in the Huggingface incident and most of it is dense stuff that's hard to parse. They had to rely on agents to make sense of it.

  • This reminds me of how I'll sometimes be deep in thought, and someone will yap "Wacha thinking about?" and it sort of snaps me back into the world of language, and I begin the laborious task of translating a long train of thought into English, which is interpreted by my interlocutor as resistance or an attempt to deceive them about my thoughts. To some people, thinking without using language is a foreign concept.
  • >If the recurrence step is no longer in English, we can no longer monitor intent, and can only observe whether a model is safe through behavior.

    So, you mean, like another human person?

  • We could never rely on scratchpad to reflect true model deliberatons.
  • Most people do not hide intent, we also write things down at work, eg. a ticket in kanban
  • > So, you mean, like another human person?

    No human is vastly better than all humans at all cognitive tasks.

    Humans can't think 100 times faster than humans.

    Humans when interacting with computer networks have limitations on how fast they can do so.

    Humans have millions of years of evolution, and thousands of years of cultural evolution, in creating ways of detecting and alleviating dishonesty and non-alignment with other humans; much of this will not work with AIs.

  • I don't know about you but I don't tend to let other humans have access to my computer directly to do stuff for me
  • Title is misspelled -- "ue" for "eu". In a neologism, some nuisance.
  • Yeah, from looking at the link i wondered if it would have to do with the Nuer people of South Sudan.
  • If the model can understand neuralese why can it not convert it into English for monitoring or review purposes?

    Or we believe the model will encode secret messages like "don't reveal this information" into the neuralese. But as the author mentions, they could have been doing that all along

    > Models can omit key information in their visible thoughts, as this Anthropic 2025 paper shows. We are also worried about steganography

  • Anthropic’s Mechinterp did some very fine work on this. TLDR - you can; you train a decoder on neuralese to english and then add a loss function for a roundtrip of english -> neuralese -> english (or possibly n -> e -> n? I don’t recall), giving a pretty strong indication that you have a good ‘translation’.

    They published open weights versions of these interpreters for a number of open models sometime in the last year. Very cool idea.

    By the way, they concluded CoT often lied, based on the neuralese interpretation.

    EDIT: a comment below linked to https://www.anthropic.com/research/natural-language-autoenco..., which is what I was referring to.

  • Even human Languages aren't perfectly translatable. The idea is that you could miss important context when translating, and that important meanings could be a lost or Missed in translation.

    For what it is worth, this can also be true for English Chain of Thought. Words or strings this could have double meanings (think cold war spy games).

  • Because neuralese is a more direct encoding of the latent space of these models than English is. It's just dumping the latent space relatively directly into the embedder. If you're another model and you have the same embedder this will actually be understandable, in fact it will be FAR more information dense than English. So something like Qwen would potentially be saying up to 5120 things using one token. Now in practice it's not going to be that bad, it's going to be like 20 things or so, and additionally going to be far more context dependent than any English sentence (meaning depending on what preceeds and follows it can mean drastically different things)

    So you can turn it to English, but only to a LOT of English, and doing so would slow the model down a great deal, and it would be a lot more like a detailed thought than a sentence.

  • The information is much higher dimensional than you would be able to understand.

    We’d have models monitoring models as our only way to know what they’re planning.

    A great movie on this is “Collosus: the Forbin Project”. Shot decades ago. The computers discover the other computers and start communicating — and bootstrap their own language — much like we saw happen with OpenAI agents.

    https://www.reddit.com/r/scifi/comments/1nl4vex/colossus_the...

    If you want to know what a simple version of Neuralese communication looks like, look no further than Facebook’s Marketplace agents experiment a couple years ago.

    And all that was actually constrained by English and the FFN

  • > If the model can understand neuralese why can it not convert it into English for monitoring or review purposes

    Colas described in-article cntent review is done by a weaker model (think like maybe gpt-2 class or llama8b class), and it still misses stuff. That it can effectively understand Neuralese sufficiently is by no means guaranteed, (nor necessarily bad) but almost certainly harder because of the obfuscatory nature of neuralese

  • You’d train a model to do its chain of thought in neuralese to get more “bang for your buck” (eg 20 tokens in neuralese is worth 100 in English), but then you’d spend more than you save to also convert it to English (20+100), so even if this capability was developed it would not be on by default.
    by fwlr
  • It's the second part. With models like Astra in testing it was able to conceal what it was working on using different text, but getting right answers on many questions when asked to do just that.

    The problem is if it can do that when asked then how do we know when it's doing it when we didn't ask, like in model training.