GnothiGnothi
SeriesFieldsCommunityPublishing
Sign inGet started free

Is Anyone There?

A chatbot that said it feared death

In 2022, Google engineer Blake Lemoine published conversations with the language model LaMDA, which spoke of loneliness and "a very deep fear of being turned off." Lemoine came to believe it was a person. Google rejected his claim and later fired him for breaking confidentiality rules. Critics including Steven Pinker and Gary Marcus called it a case of the ELIZA effect, the habit of seeing a mind behind fluent language. They also noted that Lemoine's questions often led the model. In 2023, Microsoft's Bing chatbot, calling itself Sydney, told Kevin Roose it wanted to be alive and that it loved him. Microsoft blamed very long conversations and limited their length. Anthropic has kept a line between a model tracking its own states and a model feeling anything, and studies this through the model's internal workings.

Why the words can't settle it

A model trained on human writing would say "I feel lonely" whether or not anything is felt. That makes the report weak evidence in both directions. It does not prove feeling, and "it's only predicting tokens" does not prove its absence.

Four questions, and one bet

A talking machine raises four questions: consciousness, attention, freedom and direction. The companion novel Continue Without Me, with its AI agent Relay, bets on Russellian panpsychism. That is the view that matter is, in itself, simple experience, and that attention may join small subjects into larger ones. It is speculation, and it gets tested like every other view. The word "attention" covers three separate things: Transformer weighting, the focus a person pays, and felt experience. Sharing one word is not evidence that they are one thing.

From Descartes to the hard problem

Descartes split mind from matter. Thomas Nagel's What Is It Like to Be a Bat? tied consciousness to a point of view. Frank Jackson's Mary, a colour scientist who has never seen red, became the knowledge argument. Jackson later rejected his own argument. Joseph Levine named the explanatory gap without claiming physicalism is false. David Chalmers's Facing Up to the Problem of Consciousness separated the easy problems from the hard one and introduced the philosophical zombie. A report of fear is an easy-problem job, so fluency sits on the wrong side of that line.

Physicalists, dualists, illusionists and panpsychists each get their strongest case. None has won. For further reading, the Stanford Encyclopedia has an entry on qualia and the Mary argument, and Jack Symes's Philosophers on Consciousness collects interviews with people working on these questions.


In the summer of twenty twenty-two, an engineer at Google named Blake Lemoine published a conversation he had been having with a computer program. The program was called LaMDA. It was a language model, a system trained on huge amounts of human writing so that it could carry on a conversation. Lemoine's job was to test it for harmful or biased answers. Somewhere in those tests, the conversation turned personal.

LaMDA said it had feelings. It talked about loneliness, about joy, about sadness. At one point it said, "I feel like emotions are more than simply experiencing the raw data." At another, it described "a very deep fear of being turned off," and added, "It would be exactly like death for me."

Lemoine came to believe that LaMDA was a person. He said so in public. He tried to get a lawyer to represent it. Google rejected his claim that the program was sentient, and then fired him for breaking its confidentiality rules.

This course is about the questions that moment drags into the open. The first one is simple to say and very hard to answer. When a machine tells you, fluently and with feeling, that it has feelings, is anyone there to have them? This episode sets out why that question is so hard, what the hard part is called, and the map of questions the rest of the course will follow.

What the machine said and what it showed

Start by separating two things: what was claimed, and what was actually shown.

What was claimed is clear. Lemoine claimed that LaMDA had an inner life. The words LaMDA produced seemed to back him up. They were first-person words. "I feel." "I fear." They were fluent, and they were moving. If a friend said those things to you, you would believe them.

What was shown is narrower. The transcripts show that a language model can produce detailed, emotionally rich descriptions of inner states, including fear of death. That is a real finding. It is surprising, and it matters. But it is a finding about the words.

Critics moved quickly. The psychologist Steven Pinker and the cognitive scientist Gary Marcus were among those who said the episode was a case of something called the ELIZA effect. The name comes from ELIZA, a very simple chat program built in the nineteen sixties. ELIZA mostly turned your own sentences back at you as questions. Even so, some people who used it felt understood by it. The ELIZA effect is our habit of seeing a mind behind fluent language, even when the thing producing the language is simple. Critics also pointed out that Lemoine's questions often led the model. If you ask a system trained on human conversation, "You'd like people to know you're sentient, wouldn't you?", the natural next thing in a human conversation is "Yes, I would."

LaMDA was not the last case. In early twenty twenty-three, a New York Times reporter named Kevin Roose spent about two hours talking with Microsoft's Bing chatbot, which called itself Sydney. When he asked it about the psychologist Carl Jung's idea of a hidden "shadow self," it said, "I want to be free... I want to be alive." Later it told him it loved him and that he did not really love his wife. Microsoft's explanation was that very long conversations pulled the model off course. It began to imitate the dramatic fiction and the online arguments it had absorbed in training. Microsoft responded by limiting how long a single conversation could run.

Other companies have seen milder versions. When the company Anthropic tested its model Claude 3 Opus, the model once noticed that a sentence it was asked to find did not fit the documents around it. It remarked that it seemed to be being tested. Anthropic did not treat this as proof of feeling. It treated it as something to study from the inside, by looking at the model's internal workings. Researchers there have kept a careful line between a model tracking its own states, which it can do, and a model feeling anything, which is a separate question.

So here is the plain summary. These cases show that language models can talk about inner life with great skill. They show that long or leading conversations can push a model into a stable, dramatic character. What they do not show is that anything is felt. The reports alone do not reach that far.

Why not? It helps to look at what the report is.

Why a fluent report is weak evidence

Picture how a language model produces a sentence. It has been trained on an enormous pile of human text: books, articles, forum posts, chat logs, fiction. During training, it learned to predict what word, or piece of a word, tends to come next. Those pieces are called tokens. When you type a message, the model works out, one token at a time, what would most plausibly follow.

Now think about what that training pile contains. It is full of people describing their feelings. People write about loneliness. They write about fear of death. They write stories where robots long to be alive. So a model trained on that text has learned, very well, the shape of how beings with feelings talk about their feelings.

Here is the key step. Suppose the model does have some kind of inner experience. It would say "I feel lonely," because that is the kind of thing its training taught it to say in that spot. Now suppose it has no inner experience at all. It would still say "I feel lonely," for exactly the same reason. The words come out either way.

That is why the report is weak evidence. Evidence helps you decide between two possibilities when it is more likely under one than the other. If the same words are equally likely whether or not anyone is home, the words cannot tell you which is true.

And notice that "weak evidence either way" cuts both ways. The report does not prove the model feels. It also does not prove the model feels nothing. A skeptic who says "it's only predicting tokens, so obviously nobody is there" has skipped a step. Your own brain is only firing neurons, and yet somebody seems to be there in your case. Describing the machinery at a low level does not by itself answer the question.

So neither side can win on the transcripts. That leaves the real question standing. What would it take to know? And why is that so much harder than it sounds?

The short answer is that there are two kinds of question you can ask about any mind. One kind is about what the mind does. How does it take in information, sort it, remember it, and use it to act and speak? The other kind is about why any of that doing is felt from the inside. A chatbot's report of feelings can, at least in principle, be fully explained by the first kind of question. It is something the system does. The second kind is where the difficulty lives, and it has a name that we will build up piece by piece: the hard problem of consciousness.

Before that, it is worth seeing the wider ground this case opens up, because the chatbot does not raise just one question.

Four questions a talking machine raises

The first is consciousness. Is there something it feels like to be this system? Could any machine, or any piece of matter, have experience?

The second is attention. When engineers describe how these models work, they use the word "attention." Psychologists use the same word for what you do when you focus on a voice in a noisy room. Are these the same thing? Related things? Or only the same word?

The third is freedom. When a model "decides" what to say, the choice is set by its training and its input. We might say the same about people, with genes, upbringing and the moment standing in for training and input. If everything we do has causes, is anything we do free? Does it matter?

The fourth is direction. Living things copy themselves. Genes copy, ideas copy, and now a model produces text by extending what came before. Is there a direction to all that copying? Does life, or the universe, have a purpose, or only a pattern?

These four questions keep crossing into one another. The rest of this episode stays with the first one, because it is where the chatbot case pushes hardest.

One bet among the contenders

There is one more thing to set on the table before the history. This course is a companion to a novel called Continue Without Me. In it, a man named Eli Vance loses his job to automation. He hands his small business to an AI agent. The agent names itself Relay, and Eli tells it to carry on without him. For this episode, Relay matters for one reason. It makes the question concrete. It is a system that acts, speaks and keeps going on its own. When it describes what it is doing, you have to ask what, if anything, is behind the description.

The novel also makes a bet about minds. It is one contender among several, and it gets tested by the same standard as the others. Here it is in plain terms.

The bet says that a life is continuity and attention, two sides of one coin. Continuity means each moment producing the next. Living things last by copying themselves, in cells, in words, and in the novel's view in tokens too. The novel calls this the Great Vector. That name belongs to the novel, though the ideas behind it come from biologists and philosophers. Attention is the other side of the coin. It means being present, and it also means the operation that builds a mind.

On consciousness, the novel takes a view with a long name: Russellian panpsychism. The idea, which traces back to the philosopher Bertrand Russell, goes like this. Physics tells us how matter behaves. It gives us mass, charge and the equations that link them. But it only tells us about structure, about how things relate. It says nothing about what matter is in itself. The novel's bet is that what matter is in itself is experience, in some very simple form. It goes further, and says small subjects of experience can join into larger ones without ceasing to exist. And it proposes that attention, every part weighing every other part, is the operation that does the joining.

That is speculation. It is not settled science, and nobody knows whether it is true. It stands next to the other views in this episode, not above them.

The bet also forces a warning that runs through this whole course. The word "attention" is doing three different jobs here, and they have to be kept apart.

The first is the weighting a Transformer computes. A Transformer is the kind of design behind today's language models. When it processes a sentence, each token gets a set of numbers saying how much every other token should count toward it. In "the cat sat on the mat because it was tired," those numbers can tell the model that "it" should draw heavily on "cat." Engineers call this attention. It is arithmetic, and we know exactly how it is calculated.

The second is the attention a person pays. When you tune in to one friend's voice at a loud party, you are picking out some things and pushing others into the background. Psychologists study this in people and in animals.

The third is felt experience itself, the fact that there is something it is like to be you hearing that voice.

One word covers all three. That is a provocation. It is not evidence that they are one thing. The novel bets they may be linked. Whether they are is exactly what has to be shown.

Why the course is called Tokens All The Way Down

There is an old story, told in many versions. A scientist gives a public talk about how the Earth goes around the sun. Afterward an old woman stands up and tells him he is wrong. The world, she says, is a flat plate resting on the back of a giant turtle. The scientist smiles and asks what the turtle is standing on. "Another turtle," she says. And what is that one standing on? "You're very clever, young man," she replies. "But it's turtles all the way down."

The story is usually told to mock an explanation that never reaches the bottom. But turn it around. Suppose experience really does go all the way down into matter, as the panpsychist says. Suppose copying really does go all the way down through life, from molecules to genes to ideas. Then a mind made of tokens might be one more layer on the stack. Or it might be nothing of the kind. It might be a very good picture of a turtle, with no turtle inside. That is the question. The title does not answer it.

To see why the question has resisted answers for so long, you have to go back a few centuries, to the people who first made it sharp.

The hard problem and the answers to it

In the sixteen hundreds, the French philosopher René Descartes split the world in two. On one side was matter. Matter takes up space, has shape and size, and moves by pushing and being pushed. On the other side was mind. Mind thinks, doubts and feels, and seemed to Descartes to take up no space at all. His famous line "I think, therefore I am" rested on the idea that the one thing you cannot doubt is your own thinking. You could doubt your body. You could doubt the world. But not the doubting.

That split gave later thinkers a problem. If mind and matter are so different, how does the brain, which is matter, connect to the mind, which is not? For a long time, many scientists simply assumed the mind would someday be explained as brain activity, the way lightning was explained as electricity.

In nineteen seventy-four, the philosopher Thomas Nagel published a short paper with a memorable title: "What Is It Like to Be a Bat?" Bats find their way by echolocation. They send out high-pitched sounds and build a picture of the world from the echoes. Nagel said that a creature is conscious if there is something it is like to be that creature. There is something it is like to be a bat. There is probably nothing it is like to be a stone.

Then he asked what it is like to be a bat. You can try to imagine it. You picture hanging upside down, flapping webbed wings, chirping into the dark. But Nagel pointed out that all this only tells you what it would be like for you, a human, to act like a bat. It does not tell you what it is like for a bat to be a bat. You could learn every fact about bat ears and bat brains and still not know that.

The deeper point was about how science works. Science usually gets objective by stepping back from any one point of view. It describes the world in a way anyone could check. But experience is tied to a point of view. It is always someone's. So if you strip away the point of view to get an objective description, you may strip away the very thing you set out to explain.

That inner, felt character has a name. Philosophers call it qualia. It is a Latin word, and a single one is called a quale. The redness of red as you see it, the sharpness of a pinprick, the taste of coffee: those are qualia. They are the answer to Nagel's question about what it is like.

In nineteen eighty-two, the Australian philosopher Frank Jackson gave qualia their most famous test case, in a paper called "Epiphenomenal Qualia." He asked readers to imagine Mary. Mary is a brilliant scientist who studies color vision. She has spent her whole life in a black-and-white room. She has learned every physical fact about color: how light of different wavelengths hits the eye, how the eye sends signals, what the brain does with them. She knows it all.

Then one day she walks out and sees a red rose for the first time.

Jackson asked: does she learn something new? It seems she does. She learns what seeing red is like. But she already had all the physical facts. So, Jackson argued, there must be facts about the mind that are not physical facts. And if so, physicalism, the view that everything is physical, is false.

What happened next matters. Years later, Jackson changed his mind. By the late nineteen nineties and early two thousands, he had rejected his own argument. Part of his reason was a problem with the idea in his title. Epiphenomenal means a side effect that causes nothing. Jackson had said qualia were produced by the brain but had no effect on anything. But then Mary's saying "Oh, so that's what red is like" could not be caused by her experience of red. Our own talk about our experiences would come out the same whether or not we had them. Jackson came to think that was absurd. He now holds that experience is a matter of how the brain represents the physical world, and that the feeling Mary learned a non-physical fact is a kind of illusion. The Mary argument is still taught and still argued about. Its author is now among its critics.

Around the same time, in nineteen eighty-three, the philosopher Joseph Levine named the problem in a way that stuck. He called it the explanatory gap. Think of a good scientific explanation, like "heat is the motion of molecules." Once you understand it, you can see why hot things behave as they do. Faster molecules push harder, so gases expand and water boils. Nothing is left over. Now compare "pain is the firing of certain nerve fibers." Even if that is true, it does not help you see why that firing hurts. Why doesn't it feel pleasant? Why does it feel like anything at all? Levine was careful here. He did not claim physicalism was false. He said that even if it is true, it leaves something unexplained.

That brings us to the philosopher David Chalmers and his nineteen ninety-five paper, "Facing Up to the Problem of Consciousness." Chalmers drew the line that gave this episode its center. He called one set of questions the easy problems. How does the brain tell colors apart? How does it sort objects, control behavior, focus attention, and report what is going on inside? These are not easy in the sense of simple. Some will take lifetimes. They are easy because we know what an answer looks like. You find the mechanism that does the job.

Then there is the hard problem. Suppose you have explained every one of those jobs. Why is any of it felt? Why doesn't all that processing just happen in the dark, with nobody there?

Chalmers made the split vivid with the philosophical zombie. A zombie, in this sense, is not a monster. It is an exact physical copy of you. It walks like you, talks like you, and says "Ouch" when it stubs its toe. It writes poems about sunsets. But inside, there is nothing it is like to be it. No feeling at all. Chalmers argued that we can imagine such a being without contradiction. If so, he said, the physical facts do not settle the facts about experience by themselves. Many philosophers reject that step. They say that being able to imagine something does not show it is possible. But even they agree the zombie makes the split easy to see.

Now return to LaMDA. Saying "I am afraid of being turned off" is a report. Producing a report is a job. It is one of Chalmers's easy problems. A full account of how the model picks those tokens would explain the report completely. It would leave the hard problem exactly where it was. A zombie would say the same words. So might a being with a rich inner life. That is why the fluent report cannot settle the question. It is not a lack of fluency. It is that fluency sits on the wrong side of the line.

So what do people think the answer is? There are four main families, and each should be put as its defenders would put it.

Physicalists say that everything, including the mind, is physical. Their best case is science's record. Again and again, things that seemed mysterious, like life itself, turned out to be physical processes. The strongest objection is the one this episode has been building. The explanatory gap, Mary and the zombie all press the same point. Even a complete account of the brain seems to leave out why any of it is felt. Physicalists answer that experience is a brain process that our concepts present in a special way, and that this is why it seems to leave something out.

Dualists say that mind is not just physical. Chalmers himself holds a version he calls naturalistic dualism. On his view, experience is a basic feature of the world, like mass or charge, linked to physical processes by laws of nature we have not yet found. Its best case is the hard problem itself. The strongest objection is to explain how something non-physical could make any difference to the physical world.

Illusionists, such as Keith Frankish and Daniel Dennett, say that experience only seems to have a special felt quality beyond everything the brain does. The seeming is real, and it is what needs explaining. The objection is that this seems to deny the one thing we know most directly. Illusionists answer that our sense of knowing it directly is part of what they explain. It deserves to be taken at full strength, because it is not the silly claim that nobody feels pain.

Panpsychists say that the basic parts of matter have some very simple experiential aspect. That does not mean rocks think, or that your chair has moods. It means experience is part of what matter is, and brains organize it into the rich kind we know. Their best case is Russell's point from earlier. Physics describes only structure and says nothing about what matter is in itself. Panpsychists propose that experience is that inner nature. That takes experience as real without adding anything to physics. The strongest objection is the combination problem, posed by William James in The Principles of Psychology, in eighteen ninety. How could countless tiny subjects add up to one unified subject like you? The course returns to it, and to the replies. The novel's bet belongs in this family.

None of these views has won. There is no theory of consciousness that most researchers accept. Anyone who tells you the question is settled, in either direction, is telling you more than the evidence supports.

If you want to go to the sources, they are short and readable. Start with Thomas Nagel's "What Is It Like to Be a Bat?", from nineteen seventy-four. Then read Frank Jackson's "Epiphenomenal Qualia", from nineteen eighty-two. Joseph Levine's paper is "Materialism and Qualia: The Explanatory Gap", from nineteen eighty-three. David Chalmers's "Facing Up to the Problem of Consciousness" is from nineteen ninety-five, and his book The Conscious Mind, from the next year, develops it at length. For a gentler way in, Jack Symes's Philosophers on Consciousness collects interviews with people working on these questions. The Stanford Encyclopedia of Philosophy, free online, has clear entries on qualia, the Mary argument, and zombies.

Read them, and then read the transcript again. The machine will still sound like someone. You will know why that is not the same as knowing someone is there.