The Dangers of Using "Brain-Like Language"—A Calm Re-examination of the J-space Paper by Expert Journalists
Previously in this column, I introduced the "J-space" paper by Anthropic, which claimed to have found a place for "conscious thought" within Claude. A week later, MIT Technology Review published an article interviewing their senior editor, Will Douglas Heaven, about this research, re-examining "what this discovery actually shows and doesn't show." This time, I'd like to introduce this meta-examination of the research itself.
The Unassuming Yet Essential Difficulties of Interpretability Research
First, it's important to understand the technical difficulties unique to this field. Anthropic's research area, "mechanistic interpretability," involves peering into the complex mathematical formulas of AI models to understand why a particular output was chosen.
Heaven explains this difficulty with a striking analogy. If a medium-sized LLM (Large-Scale Language Model) were printed on paper, it would cover an area the size of a city like San Francisco. Extracting what's actually happening within a massive mathematical formula consisting of hundreds of billions of parameters in a way that humans can understand requires specialized tools that know where and how to look, and creating those tools itself requires some prior knowledge of complex mathematical formulas—a cyclical difficulty.
A Perspective on Anthropic's "Narrative Style"
Another interesting point that Heaven makes in the interview is directed at Anthropic's style of information dissemination itself. Heaven prefaces his remarks by saying, "Anthropic is known for publishing strange and difficult research," and then makes a somewhat critical comment: "The narrative that they've created a very strange technology, and that only they can decipher it, perfectly fits Anthropic's image."
In this context, the series of events cited is when Anthropic warned that its new model was "too powerful to be a global cybersecurity risk," and immediately afterward, the US government restricted access to that model (access to some of these models has since been restored). Apart from the content of the research announcement itself, the observation that the company's narrative is consistently that it is "dangerously powerful, but only we can control it" is a perspective that cannot be ignored when considering how AI research is perceived.
Caution Regarding the Use of "Brain-like Language"
More than the technical content, what I think is most important in this interview is the caution regarding the use of the phrase "brain-like." Heaven frankly states, "I don't like using that kind of language." The point is that LLM is not a brain, and this kind of language risks misleading people into thinking LLM has more human-like capabilities than it actually does, or creating inappropriate assumptions about its behavior.
Anthropic itself has also expressed a cautious stance on this point. When Heaven inquired about this point, Anthropic reportedly issued a statement to the effect that, "These analogies (with neuroscience) were useful in designing experiments. We were able to make many experimental predictions about J-space that were initially not obvious, and these predictions turned out to be correct. However, it should also be noted that there are important differences between J-space (and language models in general) and the human brain, and we are not claiming a perfect correspondence." This is a self-restrained distinction: Analogy is merely a tool for formulating hypotheses, not proof of conclusions.
The practical question: "What problems can this discovery solve?"
So, what practical use does the concept of J-space have? Anthropic suggests using it as a means of capturing situations where the model is doing something it "shouldn't be doing." Since the words appearing in J-space don't appear in the model's final output, the theory is that they can provide signals of behaviors that would normally be overlooked, such as biased responses or weighing the pros and cons of dishonest behavior.
However, Heaven offers a cautious assessment here as well. He states, "This is purely theoretical," and comments, "This result should be viewed not as a practical tool in itself, but rather as another stepping stone in the journey of understanding this technology as a whole." His consistent stance is that before discussing flashy application possibilities, we must first accurately define its position as fundamental research.
What Researchers Can Learn from This Meta-Verification
This article itself does not report new experimental results, but rather is a critical analysis by a specialist journalist of already published research. However, when following AI research news, this attitude of not blindly accepting published content, but also verifying the validity of corporate rhetoric and metaphors is, though understated, extremely important.
Especially when applying terms referring to the inner workings of the human mind, such as "consciousness," "brain," and "thought," to AI models, it is crucial to always be mindful of whether those terms accurately represent the experimental results or are merely rhetorical devices designed to excessively stimulate the reader's imagination. I believe that research in this field can only progress healthily if we have both Anthropic's own restrained stance—that "analogy is a tool for hypothesis building, not proof"—and processes like this one, where external expert journalists re-examine his work.