Markov, 1913: the mathematician who read Pushkin letter by letter, long before LLMs
Before Shannon, before computers, a Russian mathematician counted the vowels and consonants of a novel in verse with a pencil. He wanted to win a theoretical quarrel. He asked the question language models still ask at every word.

When someone tells me an LLM “predicts the next word”, I think of a man who never saw a computer. In January 1913, in St. Petersburg, the mathematician Andrey Markov presented an unusual piece of work to the Academy of Sciences: with a pencil, he had counted the vowels and consonants in the first 20,000 letters of “Eugene Onegin”, Pushkin’s novel in verse.
That work says nothing about Pushkin’s poetry. It says something far more general: in a text, what comes next depends on what came before. A hundred and thirteen years later, that is still the idea that runs language models.
A quarrel between mathematicians
Markov was not trying to understand literature. He wanted to win an argument. According to science writer Brian Hayes, in American Scientist, his opponent was Pavel Nekrasov, a Moscow mathematician who had argued in 1902 that the law of large numbers applies only to independent events. Nekrasov drew a theological conclusion from it: if social statistics obey that law, human acts must be independent, and therefore free.
Markov attacked the reasoning, not the theology. From 1906 he showed that the law of large numbers also holds for linked events, provided they form what would later be called a chain: a sequence in which the probability of the next step depends on the present one. What he lacked was a real example. For years he was openly indifferent to applications. He wrote to his colleague Chuprov that he was “concerned only with questions of pure analysis”. In 1913 he changed his mind and picked a text every Russian schoolchild knows.
What Markov did, step by step
His lecture of 23 January 1913 was translated into English and published in 2006 in Science in Context (translation by Nitussov, Voropai, Custance and Link). The method is described precisely:
- The corpus. The first 20,000 letters of the novel, which is the whole first chapter and sixteen stanzas of the second, excluding the hard and soft signs, which are not pronounced on their own.
- The layout. The letters, copied without spaces or punctuation, were arranged in 200 squares of ten rows by ten columns.
- The count. 8,638 vowels and 11,362 consonants, so a vowel share of 0.432.
- The pairs. Markov counted how often a vowel follows a vowel: 1,104 times. From that he deduced 3,827 consonant pairs without counting them one by one.
Brian Hayes, who repeated part of the exercise on an English translation, estimates that Markov must have spent several days on it. Hayes himself missed 62 of the 248 vowel pairs in his sample. Markov later applied the same method to 100,000 characters of a memoir by Sergei Aksakov.
The result: the next letter is not pure chance
If letters were independent, a vowel would follow a vowel with the same probability as anywhere else, about 43%. Hayes calculates that you would then expect about 3,731 vowel pairs. Markov found 1,104, more than three times fewer.
His two key numbers fit on one line. After a vowel, the next letter is a vowel 12.8% of the time. After a consonant, 66.3% of the time. His conclusion, in the English translation:
As we can see, the probability of a letter being a vowel changes considerably depending upon which letter – vowel or consonant – precedes it.
In other words, knowing the current letter changes the forecast for the next one. A text is not a series of coin tosses. It is a series of linked states, with probabilities of moving from one state to the next. That is the definition of a Markov chain.
From Pushkin to LLMs
The link to today’s AI runs through Claude Shannon. In 1948, in “A Mathematical Theory of Communication”, he generated artificial sentences by choosing each letter, then each word, based on what came before. He noted that such processes are known mathematically as “discrete Markoff processes”. The spelling changed; the idea did not.
A modern language model performs the same operation at a different scale. Markov looked one letter back and knew only two states, vowel or consonant. An LLM looks thousands of words back and chooses from a whole vocabulary of word fragments. But the question asked at each step is the same: given what came before, what most likely comes next?
I find this lineage useful to keep in mind, because it is a reminder of what a language model is and what it is not. It has learned what usually comes next. It has not learned what is true.
What it changes for an organisation
In my work, this 1913 lesson translates into three concrete decisions.
Pick tasks where “usually” is enough
A model that knows what usually comes next is reliable on repetitive, well-documented tasks: answering frequent questions, sorting incoming requests, drafting the first version of a quote or a meeting summary. The more the task resembles what it has already seen, and the more of your own documents it is given, the better it performs.
Put a human on the exceptions
The same model is weak on anything out of the ordinary: a dispute, an unhappy customer, a case outside the procedure. There, the most likely continuation is not the right answer. Decide in advance who takes over, on what criteria, and how escalation is triggered.
Measure, then name a business owner
Markov did not claim that letters were linked: he counted it. The same discipline applies to an AI project. A rate of correct answers, a volume handled, a handling time, measured before and after. And one person, on the business side, accountable for the result. That is often what is missing when a project stays stuck at the prototype stage, as I explain in why AI pilots fail to reach production.
What I take from it
Markov spent days on 20,000 letters to win a theoretical quarrel. He never tried to make a machine write. But he asked the right question, the one language models still ask at every word they produce. The most likely answer is often the right one. An organisation’s job is to know when it is not.
In the same series
The other episodes of “the ancestors of LLMs”, in chronological order:
- Ada Lovelace, 1843: the engine “originates nothing”
- Shannon, 1948: a book opened at random
- Turing, 1950: the test, and the wrong question
- Dartmouth, 1956: the first underestimated AI quote
- Rosenblatt, 1958: the Perceptron’s promise
- ELIZA, 1966: a machine that understands nothing
- Jelinek, 1980s: statistics beat grammar
- Asimov, 1991: two intelligences, not one
Frequently asked questions
What did Markov do with Eugene Onegin?
In 1913 he analysed the first 20,000 letters of Pushkin’s novel, without spaces or punctuation. He counted 8,638 vowels and 11,362 consonants, then the letter pairs. After a vowel, the next letter is a vowel 12.8% of the time; after a consonant, 66.3% of the time.
How is a Markov chain related to an LLM?
A Markov chain predicts the next state from the current one. An LLM predicts the next word from the words before it. In 1948 Shannon already described his generated text as “Markoff processes”. The difference is scale: one letter of context for Markov, thousands of words for an LLM.
What does this mean for a company using AI?
A language model knows what usually comes next, not what is true. It is reliable on repetitive, well-documented tasks. Exceptions need a named human owner, escalation criteria and a measured result.