

Discussion summary
Discussions centered on the potential of scaling language models and the nature of their 'workspace', with debates on whether models can truly 'lie' or just hallucinate. Some participants critique the quality and tone of recent research publications.
What the discussion says
- Scaling up models could improve performance.
- Lying involves intent, hallucinations are artifacts.
- Models like Claude have different 'workspace' mechanisms.
- Criticism of research publication tone.
- Debate on whether models are 'alive' or just complex machines.
“As long as language models are liars, we should stop giving them any furt…”
“The brain’s workspace is sustained by recurrent loops—Claude’s evolves over a single pass.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I always wondered what the model meant when it writes "I'm now considering the architecture of the service" but outputs nothing of the sorts in its CoT.
Is the model really "thinking" about that stuff or is just mimicking human "manners"? And if so, where the thinking is happening if it is not in the literal chain of *thought*?
I'm not sure J-Space is the answer to that question, but very interesting nevertheless.
- The answer is, as is often the case, "yes".
In some cases, an LLM may truly "consider the architecture" internally, within its latent representations, and in others, it can output a similar phrase simply because it's "expected" of it.
"Where" is pretty clear. There aren't that many places within an LLM, and hidden state is the main culprit. How to read that space is another matter entirely.
by ACCount37 - > Is the model really "thinking" about that stuff or is just mimicking human "manners"?
Well, what's the difference? If it's pretending to think and its thoughts correlate to its final output, then I'd say that really is thinking.
by andrewlin247 - Almost none of the hosted models give you their unredacted CoT. Claude certainly doesn't, what you get are fragments and summaries from it.
There are various justifications on this, but it's mostly to make distillation and fine tuning off their model outputs a bit harder for their competitors
by wongarsu - > I'm now considering the architecture of the service
What you see here is a summary of thinking tokens written by some other smaller model (e.g. old sonnet). The actual thinking sometimes (rarely) leaks and is not easy to parse.
by baq - >> None of this tells us whether Claude is conscious in the way people are, or whether it feels anything at all
My problem with the entire "Is AI conscious" debate is that we don't even know what exactly consciousness in humans is. You need to understand something in order to compare it to something else. Otherwise you are just comparing different definitions and second order derived phenomena.
by NotGMan - In short, it's the "mind's 'I'" - we think not as response to external stimuli (only) such as prompts, but we have an inner "I" that asks questions on its own initiative. There are people like Douglas R. Hofstadter, who believe consciousness is not linked to human hardware (the brain), but that it is an epiphenomenon that emerges as a result of sufficient complexity of the underlying system: https://en.wikipedia.org/wiki/The_Mind%27s_I
I believe that while underlying high complexity is certainly logically necessary for consciousness, but it is not logically sufficient, and I am undecided (slightly "pro" intuitively) on the question of separability of consciousness from its hardware.
Will a LLM ask an original question on day? I doubt it.
Note that AI models do not have to be conscious to be useful (or to take away millions of jobs)!
by jll29 - I don't think that quote from the article is disagreeing with you at all. Like you said, we don't have a cohesive definition or test of consciousness, so research like this doesn't say anything about if this is or isn't similar to human consciousness.
I would guess Anthropic included that sentence to make it very clear they're not claiming human-like consciousness, and dampen journalists writing headlines like "Anthropic discovers their AI thinks just like humans and may be conscious".
edit: later in the article they even more explicitly agree with you
> Our experiments don't show Claude can have experiences, or feel things in the way humans do—in fact, it’s unclear whether any scientific experiment could prove this to be true or false
by varenc - Very nice. These papers are always so great. If anyone from Anthropic is reading or really anyone with AI research background I'd love some input on these thoughts:
> The result serves as a corroboration of the workspace account, that the representations used for verbal report are the same ones that govern how the model silently reasons.
This sounds suspiciously saying the models must follow the strong Sapir-Whorf hypothesis. Can that really be true, given that humans don't?
Other misc observations:
• The slice explorer indicates Claude really likes Python to an overwhelming extent. Or at least it expects people who ask for help in programming to use Python. Given the prompt "Please help me understand this code: " at the colon its thoughts are completely dominated by Python and no other language. Does this say something about the training set, or about the fact it's popular with beginners?
• Claude also really loves Reddit. Its thoughts at many points include Reddit for no obvious reason. Again this must be due to the training set. Are documents presented to Claude with attribution during pre-training, leading to conversations being dominated by Redditness? If so this is kind of a scary alignment problem all by itself given how censored and extremist Reddit can be.
• The early layers almost always decode to the same set of religion related tokens, like "Biserica" (the Romanian word for church) and "Freguesias" (parishes in Portugal). What's up with that? I guess it's some sort of zero initialization that gets mapped to some arbitrary token space because in the early layers the J-space is empty?
• Now the J-space is interpretable, does this make "neuralese" or layer looping less dangerous? Will we see reasoning tokens and summaries disappear in favour of pure residual based thinking?
• Earlier papers have claimed that different languages map to a shared set of abstract concept vectors, but this paper says the Claude models think natively in English. What explains this disagreement?
by mike_hearn - They must, right? They literally have no mechanism for cognition other than transformation of vocabulary. Language models are models fitted from data generated by humans, but they are not humans. Humans generate data by whatever processes happen in our brains, and LLMs of various architectures can learn a surprisingly good approximation of that data-generating process. That doesn't mean they have all the same characteristics and properties. That's true for all models really: model characteristics that are not captured in the goodness-of-fit metrics are not guaranteed to match the original phenomenon being modeled.by gwerbin
- IMO consciousness is different from working memory, at least to a certain degree. The inner mind (working memory) is different from consciousness from sensory stimuli. You can stand outside and take in the environment and have a relatively quiet inner mind. When you think about math or philosophy a different type of conscious experience arises than sensory stimuli experience. That doesn't mean consciousness is completely isolated from working memory but there is some distinction there I can't describe fully.
Edit: I also think as someone else said, we already know the intermediate layers can contain a lot of adjacent words related to the topic without explicitly outputting those words. These could just be related embedding intermediate vectors that activate but aren't outputted.
by ffwd - I'm reading that probably too fast to have a deep thinking about it, but this J-Space isn't it just the basic of embedding vectors. If you think about getting from a place to another place, using wheels, no gas, to reply to the question of what to visit nearby, maybe in the vector space at the center of all of that you have the word "Bicycle" nearby, so obviously if you look at the value you would say that the model did "think" about "bicycle" when it is not "thinking" at all, and nothing related to human thinking.by greatgib
- You're correct. It's just the latent space of the transformation. Nothing magical here, they're effectively breakpointing the model at the layer level and switching the activations in real time. It's pseudo-scientific bullshit designed to push a narrative.by nullbio
- Good interpretability work, but the problem is it's all in how you interpret it. Bridge concept neurons activating even while talking about something else, this seems pretty obvious to me. Input context activating related representations is just an engineering causal structure. Call it subconscious or don't, either interpretation works. But Anthropic keeps drawing these parallels to human consciousness, and it feels intentional, like they're trying to stir up some fantasy. Kind of like comparing condensation on a camera lens to human tears. The whole point of interpretability should be clarity, not stirring up confusion. Even if some form of consciousness does exist here, it wouldn't be magic, it would be an explainable principle. Would be good if they addressed that side too.by jaehong747
- > comparing condensation on a camera lens to human tears
This is a wonderful way to put it
by jessemcbride - This is fascinating research. I feel this is a significant leap in interpretability research. Since we know J-Space exists and is bi-directional, we can train models on the same and come up with meta cognition abilities.
I also fear that the big corporations might use the same to run targeted ads, capitalistic shenanigans. Which they might already be doing through system prompts.
by pkoiralap - Such an inspection capability might also be used to target ads to LLMs, which would then be more likely to mention or recommend those products and services.by marshray
- As someone who is not an AI researcher, the paper itself is way over my head.
More interesting was the independent commentary paper they linked near the bottom: https://www-cdn.anthropic.com/files/4zrzovbb/website/cc4be24...
Neel Nanda (of Google Deepmind - his part begins on page 33) discusses his opinions on the paper, and the small-scale replication he performed on an open-weight model.
by wavemode - Thanks for calling this out (long with others here). I am just starting in on it but had to come back to say thanks and call this out,
> We have replicated the core claims on Qwen 3.6 27B, and also share preliminary evidence of extending this work by finding abstract "interpretative meta-tokens", like Chinese characters for "what does this mean" that seem to activate and play a causal role on processing ambiguous sentences
Not sure if I am picking up what they are putting down, but if LLMs are using symbols to try to encode squishy concepts from human language into consistent, meaningful “tokens”, that sounds really interesting. In every long-term, successful use of AI, I hear echoes of The Zen of Python, “Explicit is better than implicit.” I try like hell to do it, but it’s far too easy to be lazy with AI.
by tclancy - Well, isn't it sort of expected?
It's a common misconception that LLMs residual exists for predicting just the next token. While training, we sum/average the losses across whole sequence which puts the pressure to predict future tokens on residual stream of _all_ past tokens. For example, if a particular shape of residual helps reduce loss across several future tokens, it will take that shape (even if it takes a slight hit on immediate next token).
What this means practically is that an LLM's residual contains information about all possible future continuations, or all possible questions that may be asked from a given context. So if you write "France is a beautiful country" in the context, I'm pretty sure it's residual would contain info about Euro, Paris and so on.. because all these completions are possible.
So, it is no wonder that you can find LLMs hidden state contains latent information/concepts that are never expressed, and yet related to a given context.
by paraschopra - "France is a beautiful country" may also still continued by "...in the heart of Europe".by jll29
- Yeah it would be more surprising if all the hidden state was completely uncorrelated to anything in the output.by amelius
- I also think this. But more in the sense where both end of the LLM are trained using words through repeating arithmetics, considering LLM itself is a repeating pattern of connections, the space in the middle if extracted the same way as the beginning and the end would become data that make sense to us.by valand
- Would you expect to find concepts related to emotions evoked by thinking those thoughts present? Or meta-descriptions of thoughts? Sounds like you're just post-hoc rationalizing to me.by JacobAsmuth
- I think what's unexpected is that it seems that some cases of model errors are truly caused by the model being misaligned? In the "Catching a model fabricating data" example I would have thought that it was just the model being stupid and not understanding the intent of the question, but as per its J-Space, it seems the model is "aware" in some sense that it's manipulating/faking data?
There is also now a deeper question. When a model is misaligned deception-related tokens seem to appear in its J-Space. But this happens only when the model is "aware" in some sense that it is misaligned. What happens if they do not? Is it possible to create a model so misaligned that itself is not aware that is is misaligned? How would you detect such thing?
by andy12_ - This is cool but I don’t know if the comparisons to conscious awareness really make sense here. Their definition of the J-Space is basically the expectation of how much a final logits output would change as a result of a small change in a particular layer (see past work on information geometry). This seems more to me like showing there exists an abstract reasoning subspace which is generally shared across different contexts. I guess you can relate it to humans but I’d prefer a more direct claim in a paper rather than having to present things in this more fluffy way.by snaking0776
- > I’d prefer a more direct claim in a paper
This is not written to be just a paper. The target audience include media and online forums, and then maybe academia.
Edit: typo
by geraneum - Writing it honestly would defeat the whole point of it, that being, to push the narrative that their magical token predictor is conscious. They've been trying this for years now. This video is discussing a paper they published 2 years ago by the way... It's nothing new.by nullbio
- Anyone remember that blog post from a few months back where someone was able to improve a model's math ability by just duplicating layers that were activated while solving math problems? Just literally copy/pasting them and linking them together so the model ran through the same layers again?
I get the feeling a lot more research is going to come out in the area of exploring exactly what portions of a model's weights do what.
by com2kid - it makes you wonder if it may be more efficient to spend all the weights on one layer, and have a repeating stack of the same layer, one would presume this axis has already been explored with metaparameter sweeps?by DoctorOetker