Idiom Insider The origins, properly sourced
idiominsider.com How we source & date Folk etymologies corrected Established 2026

The Record

Google Ngram: what it shows and what it can’t

The Ngram Viewer is a superb tool for tracing how common a word became. It is a poor tool for deciding when a word was born — and worse for proving where it came from.

Since its launch in 2010, the Google Books Ngram Viewer has become the first tool many people reach for when they want to know something about a word’s history. Type in a word or a short phrase, and it draws a graph of how frequently that string appears, year by year, across a corpus of millions of digitised books. It is genuinely useful and genuinely seductive, and the two qualities are related: the confident line it produces can invite conclusions the data cannot actually support. It is worth being precise about what an Ngram graph shows, and about the several things it cannot.

What it actually measures

The Ngram Viewer plots relative frequency. For each year, it takes the number of times your search string appears in that year’s books and divides by the total number of words published that year, so the vertical axis is a proportion, not a raw count. This normalisation is what makes the tool powerful: it lets you compare the changing prominence of a word across periods when vastly different numbers of books were printed. The horizontal axis is the year of publication of the books in the corpus. The most recent versions of the corpus, rebuilt in 2019, run from 1500 to 2019, though the early centuries are represented by far less material than the modern ones.

Used for its proper purpose, this is illuminating. If you want to watch a word rise, plateau or fall in general use over the nineteenth and twentieth centuries, the Ngram Viewer is one of the best instruments available. It can show you when a term went from rare to common, when a spelling variant overtook its rival, or when one phrasing displaced another. Those are questions about frequency over time, and frequency over time is exactly what it measures.

What it cannot tell you: when a word was born

The most common misuse is treating the point where a line lifts off the axis as the date a word was coined. It is not. An Ngram graph is not an attestation search; it is a frequency chart, and a word can exist for decades before it becomes common enough to register as a visible proportion of all published words. The line lifting off tells you when a word became frequent, which is often long after it was first used. For questions of first use, a dated-citation source — a historical dictionary, a newspaper archive — is the right tool, and the Ngram Viewer is the wrong one. The two answer different questions, and confusing “when did this become common” with “when did this begin” is the single biggest error people make with the tool.

The early end of the corpus is especially treacherous for dating. Before 1800 the amount of material is thin, so a single book can spike a line dramatically, and the smoothing that the viewer applies by default blurs a value across neighbouring years, softening exactly the sharp onsets that a dating question cares about. A wobble in 1650 may reflect one reprinted volume, not a shift in the language.

The data problems underneath the line

Even as a frequency measure, the graph rests on data with known distortions, and a careful reader keeps them in mind.

OCR errors. The corpus is built by machine-reading scanned pages, and the machine makes systematic mistakes. The most notorious is the long “s” (ſ) used in older printing, which optical character recognition frequently reads as “f”, so that historical “suck” becomes “fuck”, “sea” becomes “fea”, and countless real words are undercounted while phantoms are invented. This is why frequencies before roughly 1800 are treated with particular caution, and why an unexpected result in the early data should always be checked against actual page images rather than trusted on its own.

Corpus composition. The Ngram corpus is books, and only books — periodicals are excluded — chosen partly for the quality of their scanning and metadata. It is therefore not a neutral sample of how the language was used, but a sample of what was printed and bound and later digitised. Researchers analysing the corpus have shown that its makeup shifts over time, with scientific and technical publishing coming to loom large in the twentieth-century data, which can pull word frequencies in directions that reflect the publishing record rather than everyday usage.

No sense of reach. The tool counts appearances in books, not readers. A word that appears once in each of a thousand obscure volumes and a word that appears in one wildly popular bestseller are flattened onto the same axis; the graph cannot tell you whether a text was read by millions or by nobody. Frequency in print is not the same as currency in the culture, though it is easy to read the line as if it were.

Metadata slips. Because the whole thing is keyed to publication dates drawn from catalogue metadata, errors in that metadata — a reprint dated to its original year, or a misattributed edition — can place words in the wrong period entirely, producing citations that look early but are not.

What it definitely cannot do: prove an origin

Beyond dating, there is a deeper limit worth stating plainly: the Ngram Viewer cannot tell you where a word came from. It is blind to meaning, to context and to derivation. It can show that “posh” became common in the twentieth century, but it can say nothing about whether the word is an acronym, a borrowing from Romani, or something else, because it does not read the sentences it counts. An origin is a claim about descent and meaning; a frequency graph is a claim about how often a string of characters appears. No amount of curve-watching bridges that gap. Etymology is settled by dated, interpreted evidence — who used the word, in what sense, in what source — and the Ngram Viewer supplies none of those things.

Using it well

Two small habits sharpen its results. Turn the smoothing down to zero when you care about the exact shape and timing of a change, since the default smoothing averages each year with its neighbours and can invent gentle slopes where the real data is spiky. And use the case-sensitive and part-of-speech options deliberately: a search folding all capitalisations together will merge a common noun with a proper name, and the tagged corpus can separate, say, a verb from an identically spelled noun. These do not fix the corpus’s deeper biases, but they stop you reading artefacts of the default settings as facts about the language.

None of this is a case against the tool; it is a case for using it for what it is good at. Reach for the Ngram Viewer to trace how the prominence of a word or phrase changed over the modern centuries, to compare rivals, or to watch a usage rise and fade. Treat its curves as evidence about frequency, cross-check anything surprising against the underlying books, and be especially sceptical of the pre-1800 data. But do not ask it when a word was born, and never ask it where a word came from. For those questions it is not merely imprecise; it is answering a different question entirely, and mistaking one of its confident lines for an etymology is how a good tool ends up feeding a bad myth.