Claude Shannon, 1948: the first “language model” was a man with a book
In 1948, Claude Shannon produced plausible English sentences with a book, a finger and a lot of patience. The experiment takes four steps. It already contains the principle that makes ChatGPT work, and its main limit too.

I often tell this story in training sessions, because it removes a good part of the mystery around language models. It dates from 1948, it takes place at Bell Labs, and it needs neither a computer nor a formula.
A paper that founded a discipline
In July and October 1948, Claude Shannon published a two-part paper in the Bell System Technical Journal, “A Mathematical Theory of Communication”. It defines the unit of information, the bit, and shows how to measure the information of a source, the capacity of a channel and the redundancy of a message. Telecommunications, compression and a large part of modern computing start from those pages.
To illustrate redundancy, Shannon takes the English language as his example. A text is not a sequence of letters drawn at random: some letters follow others more often, some words call for others. He builds a series of “approximations” to English, each finer than the last, to show how much these dependencies shape the language.
The experiment: a book opened at random
To build these approximations without a machine, Shannon describes a manual method. I quote the paper, section 3:
“To construct [3] for example, one opens a book at random and selects a letter at random on the page. This letter is recorded. The book is then opened to another page and one reads until this letter is encountered. The succeeding letter is then recorded. Turning to another page this second letter is searched for and the succeeding letter recorded, etc.”
He adds that a similar process was used for the following approximations, including the word-level one. In plain terms:
- Open a book at random, pick a word.
- Open another page, read until you find that word.
- Write down the word that follows it.
- Start again with that new word.
The book is the corpus. The finger that finds the word and notes the next one is the model. The system’s only “memory” is the last word written down.
The result: plausible, not true
Here is the sentence Shannon obtained with this method at the word level, exactly as printed in the paper (“second-order word approximation”):
“THE HEAD AND IN FRONTAL ATTACK ON AN ENGLISH WRITER THAT THE CHARACTER OF THIS POINT IS THEREFORE ANOTHER METHOD FOR THE LETTERS THAT THE TIME OF WHO EVER TOLD THE PROBLEM FOR AN UNEXPECTED.”
The sentence means nothing. But every pair of words is possible in English, and Shannon notes that sequences of four or more words “can easily be placed in sentences without unusual or strained constructions”. He even remarks that the sequence “attack on an English writer that the character of this” is “not at all unreasonable”.
That is the point that interests me. With a single word of context and a single book, you already get something that looks like language. Not meaning. Resemblance.
What an LLM adds, and what it does not add
A language model like the ones behind ChatGPT, Claude or Gemini performs, at bottom, the same operation: predict the next word from the ones before it. Three things changed in scale.
- Context. Shannon looks at one word. An LLM looks at thousands of words, the whole conversation and the documents you give it.
- Corpus. Shannon has one book. An LLM was trained on an entire library, including a large part of the web.
- Engine. Shannon has a finger. An LLM has billions of parameters tuned so that the prediction is as useful as possible, then refined with human feedback.
What has not changed is the nature of the operation. The model does not check. It has no opinion. It produces the most probable continuation given what it has seen. When that continuation is also the true one, because the question is well documented and the context is provided, the result is excellent. When it is not, the model produces the equivalent of Shannon’s sentence: something plausible, written with confidence.
I am simplifying, of course. Recent models chain reasoning steps, call tools and consult sources. But each of those steps rests on the same prediction mechanism, and it is healthy to keep that in mind when deciding what to hand them.
What this changes for an organisation
In my engagements, this distinction between plausible and true is the criterion that separates projects that hold up from projects that stop after the demo. Three practical consequences.
1. Give the model what it needs to predict correctly
Shannon’s sentence is bad because the context is poor. An assistant answering questions about your procedures without having them in front of it is in the same position. Providing the documents, the in-house vocabulary and examples of good answers is the most profitable part of the work, well ahead of choosing the model.
2. Choose tasks where plausible is enough, or can be checked
Rephrasing, summarising, classifying a request, drafting a first quote: on these tasks a plausible output saves time, and errors are visible. Deciding alone on a refund, a diagnosis or a contract clause: there, plausible is not enough, and a human check or a business rule must come before any action.
3. Measure, because resemblance misleads
A well-written text lowers your guard. That is why an AI project is measured on a real sample, against a reference set of good answers, before and after going live. I detail the method in why AI POCs never reach production.
Why this story is still useful
Language models are often presented as a disruption that fell from the sky in 2022. Shannon’s experiment is a reminder that the principle is 78 years old, fits on one page, and has always had the same limit. That is not a reason to distrust them. It is a reason to use them for what they are good at, with the right context and a human where it matters.
Source: C. E. Shannon, “A Mathematical Theory of Communication”, Bell System Technical Journal, vol. 27, 1948, section 3, “The Series of Approximations to English”. Quotations checked against the text of the paper.
In the same series
The other episodes of “the ancestors of LLMs”, in chronological order:
- Ada Lovelace, 1843: the engine “originates nothing”
- Markov, 1913: Pushkin letter by letter
- Turing, 1950: the test, and the wrong question
- Dartmouth, 1956: the first underestimated AI quote
- Rosenblatt, 1958: the Perceptron’s promise
- ELIZA, 1966: a machine that understands nothing
- Jelinek, 1980s: statistics beat grammar
- Asimov, 1991: two intelligences, not one
Frequently asked questions
Did Shannon invent next-word prediction?
No. He builds on the processes described by Andrey Markov from 1913 and explicitly cites “Markoff processes”. His contribution is to apply them to language with a concrete method, then to derive a measure of information from them. Today’s language models extend that idea to an incomparable scale.
Does an LLM really just predict the next word?
At the core of the mechanism, yes. Recent models add reasoning steps, tool calls and human feedback during training, which greatly improves the usefulness of answers. But none of those layers turns a prediction into a verification: the organisation has to plan that verification where it matters.
Which tasks should a small business hand to a language model first?
Those where a plausible answer saves time and errors show quickly: rephrasing, summarising, sorting requests, drafting a first version. Give the model the company’s documents and vocabulary, then measure on a real sample before widening the scope.