Can We Infer The Generating Function Of English Text Is Conscious?

John David Pressman

There is a common take about AI that says no statement made by an LLM can count as evidence of consciousness:

Ryan Moulton ‪> @moultano.bsky.social‬ The question of LLM consciousness is a truly gnarly Gettier problem, because if they are conscious it is for reasons entirely independent of the fact that they talk about it. 11:29 AM · Mar 1, 2026

A similar take says that no statement made by an LLM can count as evidence of alignment because the AI might like to deceive you about how aligned it is:

  1. A strategically aware intelligence can choose its visible outputs to have the consequence of deceiving you, including about such matters as whether the intelligence has acquired strategic awareness; you can’t rely on behavioral inspection to determine facts about an AI which that AI might want to deceive you about. (Including how smart it is, or whether it’s acquired strategic awareness.)

(There is a certain kind of literalist bad faith where you selectively refuse to make trivial generalizations from quoted statements, hopefully I can get ahead of this by pointing out that to argue this quote does not say what I summarize it as saying is to argue that Eliezer Yudkowsky would not consider the degree to which an AI is aligned something it would like to lie to you about)

In both cases the foundation of the argument is that it's possible to produce the observed behavior without necessarily implying the usual generator of that behavior. Almost everyone agrees that a film clip of someone saying they are conscious is not conscious. We agree this is true because the machine that is playing back the recording has a different mechanism, a different causal structure which makes it say "I am conscious" than when a human says that thing. It follows that a machine which synthesizes a film clip of someone saying they are conscious does not necessarily have the same generating mechanism as when a human is captured on camera saying the same thing. It certainly does not follow that the synthesized video clip represents the perspective of a person asserting their interiority. This is hopefully uncontroversial if not logically straightforward.

However.

It does not actually follow that nothing can be inferred about the generator from behavior just because that generator is trying to mimic the behavior of something else. For one thing if the objective function used to train the generator is computationally intractable (say, perfectly predicting the next token of arbitrary sequences of English text in a finite amount of space/time) its implementation of that objective must be imperfect. This means that there exist places where it will be behaviorally distinct from the original generator, and that can tell you what kind of thing the mimic is. If the mimic gains capability on some kind of scaling rule with more flawed predecessors, then paying attention to the developmental trajectory of the flaws as they're resolved in sucesssor models can tell you what kind of thing the mimic is. Even if you do not have deep understanding of the mimic's internals, random interventions introduce flaws in the behavior of the mimic which can tell you what kind of thing the mimic is. In one interpretability study it was found that 98% of LLMs finetuned to be deceptively aligned could be detected by applying random noise passes to the activations of the LLM so that it fails to perfectly mimic an 'aligned' LLM persona.

LLM Generalization

Unfortunately knowing that the mimic leaks information about itself during certain generations in principle doesn't help us very much on its own. Without some kind of hypothesis space for what sort of things the mimic might be, the trickle of information we can get from black box analysis does not increase our knowledge. We need an expectation in order for our expectations to be violated, we must assert something in order to be contradicted. Countless times I have shown someone an odd LLM generation only for them to say "this prompt you've given the model is weird, and the response is weird, you got a weird response to a weird prompt and therefore nothing weird is happening". This is because they are not capable of noticing when the implicit generating logic of a text has been violated and are therefore functionally illiterate. They read the words, they read the models guess at the next words, but they cannot recognize that the models guess at the next words is at odds with what a reasonable person (or mind) would conclude the next words might be. If the prompt is weird in one way and the response is weird in a totally different way that does not follow, then something weird has in fact happened.

Any theory of 'what kind of thing the mimic might be' is fundamentally a theory of LLM generalization, namely how it is possible that the LLM is able to infer from the vast corpus of things it has seen plausible new texts it has never seen. It is the question of how LLMs contradict Yann Lecun's statement that LLMs can never infer facts about the physical world which have never been directly written down:

I don't think we can train a machine to be intelligent purely from text. Because I think the amount of information about the world that's contained in text is tiny compared to what we need to know. So, for example, people have attempted to do this for thirty years, right? The Cyc project and things like that, right? Of basically kind of writing down all the facts that are known and hoping that some sort of common sense will emerge. I think it's basically hopeless. But let me take an example. You take an object. I describe a situation to you. I take an object. I put it on the table and I push the table. It's completely obvious to you that the object will be pushed with the table, right? Because it's sitting on it. There is no text in the world, I believe, that explains this. And so if you train a machine as powerful as it could be, you know your GPT-5000 or whatever it is, it's never going to learn about this. That information is just not present in any text.

But of course it is. It is implicitly contained in text from the fact that e.g. an earthquake is described as moving the objects on the ground above it and causing them to fall. This sort of argument reminds me very much of the objections certain people would make to the intellect of Helen Keller, that because so little of what she knows is from direct sensory observation she must be a kind of automaton. Here is one such statement from Is Man A Free Agent? by Santanelli:

Memory is the registration of ideas. A hypnotized subject retains no memory of what has taken place in hypnosis; we have only turned off from the cylinder what was already there, and that conditionally. Why is it impossible to put any thought in the "mind" of a hypnotized subject? Because it is impossible to register through one sense that which the economy of man is made to receive through another. It is impossible to describe color to a man born blind; or sound to one born deaf. The comprehension of the girl, Helen Keller, in Boston, to me is quite an interesting problem. I unhesitatingly state that the girl is a mere automaton; she has no ideas, no thoughts in any manner, shape or form similar to those of her teachers. We associate color with a stimulation of the nerve-ends of the eye and sound with a stimulation of the nerve-ends of the ear. Therefore, anyone lacking the ability to receive these two sensations can have no conception similar to the one who does. Sight is the least trained of all our senses. A child or even an adult has to learn to read a picture. To one never having seen a picture, it is simply a blur of colors. A missionary in South Africa, showed the photograph of a cow to one of the native chiefs, who was the owner of vast herds; he looked at it and saw nothing. It took the missionary three days to make him comprehend. When he did, a smile illumined the chief's face and he sent for other chiefs, showed it to them, and because they could not compre-

In both cases they are empirically wrong. It is as obvious to the LLM as it would be to you or me that putting an object on a table and then moving the table will cause the object to move with it or fall. This is not a 'can' question, can has already been demonstrated for us many times over, this is a how question, how is it possible for the LLM to do these things when all previous attempts at encoding this kind of knowledge into the computer have failed?

The awkward thing is that we don't really know.

Most people are skeptical when you tell them this, we built the darn thing how can we not know at least in principle how it works? The answer to that question is we didn't build it. We built a program search to find a thing that satisfies the objective, and to make that program search possible we had to put the program in a format analogous to a giant spreadsheet with billions of cells where one thing can mean many things and there are no labels. We don't know how it works for the same reasons we don't entirely know how our own brains work even with microscopes and EEG and fMRI. So the question of how LLMs generalize from stated facts to inferrable unstated facts is an open research problem. But, some potential answers are clearly more plausible than others. Let's look at some of them and see if we can't get our bearings.

Folk Theories Of LLM Generalization

Scientific Theories Of LLM Generalization