Frederick Jelinek and IBM: the day statistics beat grammar
In the 1970s, an IBM team stopped writing grammar rules and started estimating probabilities from data. Its leader, Frederick Jelinek, had once wanted to become a linguist. His method produced the “language model” that today’s LLMs descend from.

When someone tells me an LLM “predicts the next word”, I think of a group of IBM engineers who, in the 1970s, stopped writing grammar rules and started counting words. Their leader was Frederick Jelinek. He had once wanted to become a linguist.
He is mostly remembered for a quip about linguists. His story deserves better: it is the story of a way of working that won because it measured its own errors.
An engineer who wanted to be a linguist
Bedřich Jelínek was born in 1932 in Czechoslovakia. According to the obituary by Jan Hajič, his family lived in Kladno, near Prague, and his father died in Theresienstadt in the last days of the war. After the 1948 communist coup, his mother emigrated with him to New York. He took evening classes at City College, then went to MIT, where he wrote a thesis in information theory, the field founded by Claude Shannon. He received his doctorate in 1962.
My favourite detail comes from his own account, published in 2009 in Computational Linguistics (“The Dawn of Statistical ASR and MT”). In the early 1960s his wife Milena attended linguistics lectures at MIT, several of them taught by Noam Chomsky. Jelinek sat in and, as he puts it, “got the crazy notion” of switching from information theory to linguistics. His adviser, Robert Fano, would not hear of it. At Cornell, the linguist Charles Hockett offered to work with him, then dropped the idea to compose operas. Jelinek spent the next ten years on information theory.
1972: IBM takes on continuous speech
In 1972 he spent the summer at IBM in Yorktown Heights and took over the new Continuous Speech Recognition group. The reason for the project, he says, was mundane: IBM worried that computers would one day be powerful enough for every need, and speech recognition promised to consume a lot of computing cycles.
Management decided the group would need linguists and assigned three. When it came to modelling natural English, one of them, Stan Petrick, reassured the team: “Don’t worry, I will just make a little grammar.” According to Jelinek, he never did, and the phrase became legendary in the group.
This was a choice against the current. Mark Liberman recalls in his 2010 obituary that in 1969 John Pierce, a leading Bell Labs figure, compared the appeal of speech recognition to schemes for “turning water into gasoline”. Artificial intelligence research at the time saw intelligence as applied logic, with hand-written rules.
The method: a noisy channel and probabilities
The IBM team framed speech recognition as a transmission problem. A speaker wants to say a string of words; the acoustic signal is a garbled version of it; the machine must recover the most probable word string. It looks for the sentence that maximises the product of two probabilities, which gives three building blocks:
- The acoustic model: the probability of observing this signal if a given sentence was spoken.
- The language model: the probability, before hearing anything, that the speaker would say this sentence.
- The search: an algorithm that explores possible sentences to find the best one.
Jelinek notes that the terms “acoustic model”, “language model” and “perplexity” were invented at IBM. Around 1978, for a 5,000-word dictation system named Tangora, the team modelled English with trigrams, at John Cocke’s suggestion: the probability of each word depends on the two words before it. Those probabilities were estimated from text, not set by hand.
The quip, and what it really says
The line usually circulates as “Every time I fire a linguist, the performance of the speech recognizer goes up”. I found no primary source for that exact wording. Here is what is documented:
- In 2004, in his talk “Some of my Best Friends are Linguists” at the LREC conference, Jelinek himself put up the line “Whenever I fire a linguist our system performance improves”. He attributed it to his presentation at a workshop on the evaluation of natural language processing systems, in Wayne, Pennsylvania, in December 1988.
- The workshop report, published in 1990, confirms that he presented an evaluation method there, but does not record the line.
- Other witnesses give other versions. Roger Moore places it at a 1985 workshop, with “phonetician/linguist”. The dedication of the ACL 2011 proceedings speaks of a linguist who “leaves the group”.
The same 2004 talk gives the substance behind the joke. On IBM’s first task, an acoustic model whose statistics were estimated by experts reached 35% accuracy. The same model, with statistics estimated automatically from data, reached 75%. Jelinek adds:
My colleagues and I always hoped that linguistics will eventually allow us to strike gold
The quip was not aimed at linguistics. It was aimed at parameters set by hand, without measurement.
From speech to translation, then to LLMs
In 1987 the team applied the same formula to machine translation. French played the role of the garbled signal, English the original message. The data came from the debates of the Canadian Parliament, published in French and English. The 1990 paper by Peter Brown, Jelinek and six colleagues describes about three million sentence pairs extracted from those debates. On 73 test sentences, 48% of the translations were judged acceptable. Jelinek recounts that an earlier submission to the Coling 1988 conference had been rejected, with this reviewer comment: “The crude force of computers is not science.”
The resulting translation system was called Candide. On an internal test of 100 short sentences, its 1992 version translated 45 correctly, its 1993 version 62.
The link to LLMs is direct. A language model, in IBM’s sense, estimates the probability of the next word given the previous ones. In 2008, in Brno, Jelinek still summed it up that way in front of his slide. LLMs perform the same operation, with neural networks and thousands of words of context instead of two. Jurafsky and Martin’s reference textbook traces one of the roots of LLMs to the same lab, with the maximum entropy language models developed there in the early 1990s.
What this changes for an organisation
What strikes me is that statistics did not win on a cleverer idea. Mark Liberman says it in his obituary: Jelinek’s most lasting contribution is a way of working, comparing competing methods on quantitative criteria defined in advance, on the same training and test data. In my engagements, I draw three rules from it.
Collect real examples before choosing a tool
IBM made progress when it replaced expert estimates with data. For a small or mid-sized business, the data are the emails, customer requests, quotes and tickets actually received. Gather a representative sample, with the expected right answer for each one.
Fix the error measure before comparing
Share of correct answers, error rate, handling time: choose the measure before testing, then apply it to every option on the same examples. Without that, you compare demos, not results. I describe such a scorecard in measuring an AI support assistant.
Keep domain expertise, in the right place
Jelinek never gave up on linguists. He relied on them to build annotated resources, not to tune parameters by hand. In a business it is the same: the domain expert defines the cases, the right answers and the exceptions. The measurement then tells you which method works.
What I take away
The line about linguists survived because it is funny. The lesson lies elsewhere: a team that measures its error on real examples ends up beating a team with better ideas that does not measure. LLMs were born from that discipline. So are the AI projects that last in a business.
In the same series
The other episodes of “the ancestors of LLMs”, in chronological order:
- Ada Lovelace, 1843: the engine “originates nothing”
- Markov, 1913: Pushkin letter by letter
- Shannon, 1948: a book opened at random
- Turing, 1950: the test, and the wrong question
- Dartmouth, 1956: the first underestimated AI quote
- Rosenblatt, 1958: the Perceptron’s promise
- ELIZA, 1966: a machine that understands nothing
- Asimov, 1991: two intelligences, not one
Frequently asked questions
Did Jelinek really say “Every time I fire a linguist…”?
The exact wording is not established. In 2004 he himself showed the line “Whenever I fire a linguist our system performance improves” and tied it to a December 1988 workshop in Wayne, Pennsylvania. Other witnesses give variants and other dates.
What is a language model in IBM’s sense?
It is the probability that a speaker will say a given string of words. IBM first estimated it with trigrams: each word depends on the two before it. The term was coined in Jelinek’s group; today it names LLMs.
How is this linked to machine translation?
From 1987 the team applied the same formula to French-to-English translation, learning from the bilingual debates of the Canadian Parliament. This led to the founding 1990 paper and then to the Candide system.