History of AI · 6 min read

ELIZA, 1966: why we trust a machine that understands nothing

In 1966, an MIT program made of keywords and rewriting rules imitated a psychotherapist. Its inventor’s secretary, who knew it was a program, asked him to leave the room so she could talk to it alone. Sixty years later, the lesson applies to every AI assistant.

Portrait of Joseph Weizenbaum, archive photograph
Joseph Weizenbaum. Photo : Rochester Institute of Technology, 1982, domaine public (Wikimedia Commons)

In 1966, Joseph Weizenbaum, a researcher at MIT, published a ten-page paper in Communications of the ACM about a program called ELIZA. You typed a sentence on a typewriter connected to the computer, and the computer answered. Nothing more. Yet sixty years before ChatGPT, people were already confiding their worries to a machine.

What interests me in this story is not the technical feat. It is what it says about us, and about how an organization should decide to trust an AI assistant.

A program that plays therapist

ELIZA ran on MIT’s time-sharing system, on an IBM 7094. The program itself knows nothing: it applies a script. The best known, DOCTOR, makes the machine answer like a Rogerian psychotherapist, the kind who keeps the patient talking by rephrasing what they just said.

The choice was deliberate. Weizenbaum explains it in the paper: the psychiatric interview is one of the few conversations in which one participant is “free to assume the pose of knowing almost nothing of the real world.” If you tell a psychiatrist “I went for a long boat ride” and he replies “Tell me about boats,” you do not assume he knows nothing about boats. You assume he has a reason for asking. You fill the gap yourself.

The paper reproduces a typical conversation. The capitalized lines are the machine’s:

Men are all alike.

IN WHAT WAY

They're always bugging us about something or other.

CAN YOU THINK OF A SPECIFIC EXAMPLE

Well, my boyfriend made me come here.

YOUR BOYFRIEND MADE YOU COME HERE

One detail: none of the replies ends with a question mark. On that system, “?” deleted the line. The machine asks questions without ever being able to mark them as such, and nobody minds.

The mechanism: keywords and rewriting rules

The method comes down to four steps, described in detail in the 1966 paper:

  1. Find a keyword in the user’s sentence, in order of priority (“mother,” “I am,” “you”…).
  2. Decompose the sentence according to a pattern tied to that keyword.
  3. Reassemble it with a stock phrase, swapping pronouns: “I” becomes “YOU,” “me” becomes “you.”
  4. If no keyword is found, fall back on a content-free line such as “PLEASE GO ON.”

The original code, long thought lost, was found in 2021 by Jeff Shrager and MIT archivist Myles Crowley in Weizenbaum’s papers. His estate released it under a CC0 license. According to Shrager, the main program is about 230 lines long. No knowledge base, no lasting memory of the exchange. Weizenbaum says so himself: ELIZA throws away each of its inputs, with a few exceptions.

The secretary who asked to be left alone

Weizenbaum told what happened next ten years later, in Computer Power and Human Reason (1976). He was startled by how quickly users of DOCTOR became emotionally involved with the program. Then comes the anecdote:

“Once my secretary, who had watched me work on the program for many months and therefore surely knew it to be merely a computer program, started conversing with it. After only a few interchanges with it, she asked me to leave the room.”

She knew it was a program. When he suggested reviewing the overnight conversations, he was accused of spying on people’s most intimate thoughts.

He drew a conclusion he never let go of: “extremely short exposures to a relatively simple computer program could induce powerful delusional thinking in quite normal people.” The shock also came from his peers. Some psychiatrists imagined DOCTOR could become a nearly automatic form of psychotherapy; one text he quotes suggests “several hundred patients an hour could be handled.” Weizenbaum spent the rest of his career warning against handing judgment over to machines.

What Weizenbaum had already written in 1966

Weizenbaum is mostly remembered for his warnings of the 1970s. Rereading the original paper, I find that the essentials are already there. Two sentences in particular:

“ELIZA shows, if nothing else, how easy it is to create and maintain the illusion of understanding, hence perhaps of judgment deserving of credibility. A certain danger lurks there.”

He adds that important decisions increasingly tend to be made in response to computer output.

The second sentence is the one I quote most often with clients:

“The crucial test of understanding, as every teacher should know, is not the subject's ability to continue a conversation, but to draw valid conclusions from what he is being told.”

The link with today’s LLMs

Technically, an LLM has nothing in common with ELIZA. It does not look for keywords: it predicts the continuation of a text from billions of examples, with a fluency ELIZA never came close to. It knows things, in the sense that its training absorbed an enormous amount of information.

But the effect Weizenbaum described, now known as the ELIZA effect, is stronger, not weaker. The more fluent the conversation, the more understanding we attribute to the other side, and therefore judgment. A model that answers a question about a delivery date with confidence gives the same impression as a colleague who knows the file. Yet it has checked neither the stock nor the contract, unless someone connected it to them.

What this changes for an organization

When a team shows me an AI assistant “that works well,” I almost always ask the same question: how do you know? Too often, the answer is a demo. A convincing conversation. That is exactly the criterion Weizenbaum rejected in 1966.

Before an assistant talks to customers, I ask for three things in writing:

  1. What it must check before answering, and in which source: stock, prices, lead times, contract terms, customer history.
  2. What it is not allowed to assert: a delivery commitment, a commercial gesture, a diagnosis. In those cases, it hands over to a person.
  3. Who reviews its answers, how often, and on what sample. A named business owner, not “the AI team.”

Review is measurable: share of correct answers on a reviewed sample, escalation rate, errors reported by customers. I detail these indicators in measuring an AI support assistant. Without them, you judge your tool the way Weizenbaum’s secretary judged DOCTOR: by the feeling that it understands you.

What I take from Weizenbaum: know what a machine actually does before trusting it with a decision. Sixty years later, it is still the right question to ask of any AI project.

Primary source: J. Weizenbaum, “ELIZA—A Computer Program For the Study of Natural Language Communication Between Man and Machine”, Communications of the ACM, vol. 9, no. 1, January 1966, pp. 36–45.

In the same series

The other episodes of “the ancestors of LLMs”, in chronological order:

Frequently asked questions

What is the ELIZA effect?

It is the tendency to attribute understanding, and therefore judgment, to a program that converses plausibly. Weizenbaum observed it in 1966 with ELIZA, a program that only spotted keywords and rephrased the user’s sentences.

Did ELIZA work like an LLM?

No. ELIZA applies handwritten rules: a keyword triggers a rephrasing. An LLM predicts the continuation of a text after training on huge corpora. What they share is elsewhere: in both cases, a fluent conversation does not guarantee a correct answer.

How do you avoid the ELIZA effect with a business AI assistant?

By writing down what the assistant must check before answering, what it is not allowed to assert, and who reviews its answers, how often and on what sample. The tool is then judged on measured indicators, not on the impression left by a demo.