← Home

The J-Space is cool, but it's not consciousness

What it would take to explain subjective experience.

July 25, 2026

On July 6th, Anthropic published a new interpretability tool that they call the J-Space:

As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words. Most of this processing is invisible to you. But some of what takes place in your brain you do have access to—an image that pops into your head, or a deliberate plan you make about where to go shopping. Neuroscientists and philosophers sometimes refer to the latter type of brain activity as “consciously accessible,” to distinguish it from all the other processing that goes on unconsciously. This activity has special properties: we can describe it, control it, and use it for deliberate reasoning, in contrast to all the automatic processing that goes on without our awareness.

In a new paper, we present evidence that a similar distinction has emerged in modern language models like Claude. We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a special role.

We call the collection of these patterns the J-space—named after the technique we used to find them, involving a mathematical concept called the Jacobian.

While Anthropic has published interpretability tools in the past (notably Golden Gate Claude), this is the first paper where I've seen them describe a part of Claude's processing as analogous to conscious thought. Rather than claiming directly that Claude is conscious, the authors state that they are analyzing Claude as if it has conscious thoughts, and that they still can't tell whether Claude has subjective experience: the experience of what it is like to be something.

But what is the real difference between claiming your model behaves "as if" it has conscious thought, and claiming that it is genuinely conscious? It seems to me that the authors have overstepped by tacking on the metaphor of "consciousness" to what is simply another interesting finding in interpretability.

A similar phenomenon happened a few years ago with the co-opting of the word "hallucination" to refer to language models making up plausible-sounding but incorrect facts. This phenomenon should have been called confabulation, but unfortunately nobody consulted me before coming up with the analogy.

With consciousness, the stakes of correctly labelling the phenomenon are much higher, because calling a language model conscious implies that we should consider giving a level of moral and social consideration to machines. To see why the label is misguided, it's worth understanding what consciousness actually is, and more importantly, what it would mean for humans to truly understand consciousness. Following David Chalmers, we can divide the study of consciousness into 2-3 areas:

  1. The "easy" problem: assuming subjective experience exists, explaining the behavioral manifestations of consciousness (e.g. global workspace theory)
  2. The "hard" problem: explaining why subjective experience exists
  3. The "real" problem (as coined by Anil Seth)1: assuming subjective experience exists, explaining how specific physical processes lead to specific qualities of subjective experience.

The global workspace theory, first proposed in 1988, is one of a number of competing theories focused on the easy problem, and it's the one that inspired the J-Space research. Another example is higher-order thought, which thinks of consciousness as the ultimate type of metacognition, or thinking about thinking.

Focusing on the easy problem — the behavioral function of consciousness — skips over its most important aspect, which is subjective experience. For instance, the global workspace theory claims that consciousness helps us to focus and act on the important aspects of unconscious processes in the brain. By this definition, however, even a basic voting algorithm could be considered conscious, in that the aggregated counter variables form the "global workspace" and the operations on the underlying sequence of votes are the "unconscious." Any theory that fails to account for the nature of subjective experience is doomed to fail in this way — it's all too easy to shoehorn the theory into labelling many non-conscious systems as having a "global workspace."

As for the hard problem, the closest attempt to a unified definition of consciousness comes from integrated information theory (IIT), which states that consciousness is proportional to the amount of information a system generates over and above the information in its constituent parts. For instance, a single neuron generates the information available in its firing pattern, but group of neurons generates information both in its individual firing patterns and the sequence of firing patterns across the group. The theory makes intuitive sense, and it has found clinical applicability in the technique of PCI for measuring whether patients are conscious.

But IIT is a little crazy in its general form, because it implies that any system with integrated information is conscious. For instance, two humans having a conversation, though the level of integration is low, would generate an emergent consciousness experiencing "what it is like" to be the two people. So either we are grossly underestimating the amount of consciousness in the world, or IIT needs some refinement.

In lieu of a theory that solves the hard problem, what we can do robustly is explain how, in humans, specific physical processes in the brain lead us to have specific subjective experiences, and see if similar processes apply to artificial systems. This is what Anil Seth calls the real problem of consciousness.

The brain is a prediction machine. What that means is, rather than taking in information and composing it bottom-up into edges, shapes, and then objects, we instead form visual predictions, and then constrain those predictions based on what the eyes take in. If you believe this is how the brain works, then the nature of specific conscious experiences can be explained by specific predictions.

Consider the experience of touching a hot stove and how it related to emotions: first, your finger goes to touch the stove. You touch the stove, and before you even notice, there is a reflexive loop that forces you to pull your hand away. Meanwhile, the pain routes through a slower path to the sensorimotor cortex in your brain, and then to the areas that govern conscious experience, after which you feel the pain. The fact that you feel pain in particular, as opposed to joy or nervousness or the color green, can be explained by which particular synapses contacted the consciousness-relevant areas, as well as the specific firing pattern of those synapses.

In predictive processing terms, you started with the prediction about yourself — that your body was physically safe within its normal operating range of temperature. This was experienced as a sense of safety. You predicted your finger was going to touch the stove, and your expectation that the stove had a low chance of being hot was starkly updated by the signal of the hot stove. This prediction, in turn, was merely a signal constraining the introspective prediction of emotion, which is now revised to a feeling of shock and even anger.

More generally, the brain maintains constant high-level predictions about the condition of the body which correspond to our subjective experience of emotion. These high-level predictions are constrained by the signals we receive from sensory neurons like the ones on your fingertip. The same framework can be extended to other conscious experiences like visual perception and sense of self.

According to Seth, we can dissolve the hard problem of consciousness — the question of why subjective experience exists — by explaining how each of these varieties of conscious experience arise from physical processes in the brain. If we take these theories together, we may arrive at a mechanistic definition of subjective experience that allows us to answer whether machines are conscious.

As Seth describes in Being You, subjective experience is analogous to life. In Greek times, you could be forgiven for thinking that the wind and fire were living beings. Fire, for instances, moves without being pushed by any human, consumes fuel, can reproduce by burning more fuel, and even responds to stimuli like a wet patch of grass by skirting around it

Now that we understand biology, no one is seriously asking whether fire is living. We've defined life in a particular way that is grounded in observations, which involves a specific set of properties that includes not just reproduction, metabolism, and response to stimuli — which fire does exhibit — but also cellular organization, heredity, and homeostasis. Even though the definition of life is somewhat arbitrary, it's a convenient concept to help people choose careers, organize conferences, and explain the world on the basis of derivative terms like "biology" and "evolution."

Life, then, is not so much a property, but a label we assign to a set of systems (i.e. living organisms) that we want to group together. Consciousness, similarly, can be decoupled into aspects like sense of self, emotions, and subjective memory, but could really be defined more simply as "what it is like to be alive."

Would an AI system that fundamentally works very differently from the brain, for instance by running on totally different hardware and relying on feedforward, bottom-up perceptual approaches rather than predictive processing, know what it is like to be alive? Seth predicts that being conscious has "more to do with being alive than being intelligent." However, we cannot truly answer the question until the real problem has been explained, and it would be premature to claim otherwise.

Whether or not AI is conscious, the key underlying question is how we should treat AI practically. Do we treat AI as any other software, or do future AI agents deserve their own inalienable rights to life, libery, and happiness? Should we be okay with humans who choose to marry AIs, even knowing that they are likely not conscious?

This set of questions implies that subjective experience is a precondition for consideration in human society. But machines, far from fitting into our existing frameworks of ethics, question the framework itself.

If sentience itself is ill-defined, as machines demonstrate, what is the real line between a pig that happens to encode similar pain-response mechanisms to a human (due to evolutionary quirks) and a fire that recoils in the face of water, much less a language model whose chain-of-thought says "ouch"? Sentience, due to its incomparable and ill-defined nature, cannot possibly be a reliable basis for morality — we need more modern tools for a modern time.

The liberal consensus of the Enlightenment, and the concept of human rights, grounded as they are in the idea that humans have free will and that suffering is the primary evil to be contended with in the world, is going to be increasingly challenged by advances in our understanding of the biology of subjective experience.

No one is asking anymore whether we need to make offerings to the god of fire or the local fire spirits, because we have defined them as non-living and created a system of morality that prioritizes the living, and particularly the sentient. We should already recognize that it would be absurd to give any "human" rights to systems that are fundamentally so different from humans. But this moral shift may lead us to question the basis of human rights itself. We are going to, at some point, agree on a collective definition of consciousness that is scientifically useful. My guess is that this definition will not extend to current machines.

In defining consciousness, we will also strip it of the ambiguity that gives it universal power. We will decouple it into constituent parts that make it analyzable scientifically, and it will become clear how arbitrary it is to use this particular set of physical traits as a basis for a universal morality. Though it may take hundreds of years, we will eventually come to terms with the inconsistency of using conscious experience as a basis for morality. Whether our belief in the mythos of "sentience" supports utilitarians who want to minimize suffering or Kantians who want to insist that humans, having free will, should never be treated as a means to an end — the mosaic of moral beliefs that supports our current understanding of human rights is likely to crumble eventually.

Just as Christian ideas of the soul and good works became supplemented by a science-based ethics grounded in human rights, liberty, and the pursuit of happiness, the next moral system is likely to inherit many of the same principles of our current system, while shifting the basis for those principles in a way that allows us to make decisions about a more complex set of issues. One way this might play out is that we generalize our attitudes about humans to attitudes about systems that resist disorder in general. Living things, for instance, are especially unique in that they actively construct order in a universe intent on destroying it.

We have a fundamental bias towards anthropomorphism — to fail the Garland test and treat the robots as human even when we know they are not. We are in an era of neuroscience analogous to the early days of biology, where people were just starting to question whether fire should be considered living. The terms we assign to artificial systems, whether "hallucination" or "consciousness," can quickly become unquestioned parts of the collective narrative.

The J-Space is genuinely a powerful tool for understanding language models. But equating it to consciousness relies on the flawed assumption that the global workspace theory is a good explanation of consciousness, when in fact those explanations are waiting to be discovered.


  1. Much of the neuroscience content in this essay is drawn from Being You by Anil Seth, a book I recommend for anyone interested in consciousness. Any mistakes are my own.