Zipf's Law: Why the 2nd Word Is Half as Common

Zipf's law says that if you rank words by how often they appear, each word's frequency is roughly inversely proportional to its rank. The most common English word, "the," shows up about twice as often as the second-ranked "of," three times as often as the third-ranked "and," and so on, following a simple 1/rank curve.

That pattern is startling because nothing forces it. Writers choose words freely, yet the aggregate frequencies of millions of texts settle into nearly the same lopsided shape. The rank-2 word lands near half the rank-1 word, rank-3 near a third, rank-4 near a quarter. Count the words in any large English corpus and you will see it. You can also test it yourself at the end of this page with the Zipf's law calculator.

The 1/rank rule in plain terms

Write down every distinct word in a long text, sort them from most to least frequent, and assign rank 1 to the most common, rank 2 to the next, and so on. Zipf's observation is that frequency falls off as one over the rank. In equation form, with f the frequency and r the rank:

f(r) ≈ C / r^α     (Zipf's law, α ≈ 1)

When the exponent α equals 1, doubling the rank halves the frequency. American linguist George Kingsley Zipf estimated that for English the most frequent word occupies about one-tenth of all word tokens, giving the famous approximation f(r) ≈ 0.1 / r. So "the" is around 1 in 10 words, "of" around 1 in 20, "and" around 1 in 30. A study of 100 typologically diverse languages found the exponent clustered tightly around 1, ranging only from about 0.76 to 1.44.

The cleanest way to see Zipf's law is on a log-log plot. Take the logarithm of rank on one axis and the logarithm of frequency on the other, and the points fall along a nearly straight line with slope close to negative one. A straight line on log-log paper is the signature of a power law, and that is exactly what Zipf's law is.

A pattern with a tangled history

Zipf gets the name, but he was not first. The French stenographer Jean-Baptiste Estoup noticed the regularity in shorthand frequencies around 1916, before Zipf published his linguistic studies in 1932 and 1935. The matching rule for city sizes was published by German physicist Felix Auerbach in 1913, and mathematician Alfred Lotka gave it a modern analytical form in 1925.

Zipf's contribution was popularization and ambition. His 1949 book Human Behavior and the Principle of Least Effort: An Introduction to Human Ecology, a 573-page volume, argued that the same statistical signature ran through language, geography, and economics, all driven by humans minimizing effort. The sweep of that claim is why his name stuck to a phenomenon several others had documented first. Even the "least effort" idea predates him, having been articulated by Guillaume Ferrero in 1894.

Why the same curve shows up everywhere

The truly strange part is that the 1/rank rule escapes language entirely. Rank the cities of a country by population and the second city is often roughly half the largest, the third roughly a third, and so on, a pattern geographers call the rank-size rule. The same near-straight log-log line appears in firm sizes, where Robert Axtell's 2001 analysis in Science found U.S. company sizes following a Zipf distribution with exponent near 1, and in website traffic, documented by Adamic and Huberman in 2002, where a few sites capture enormous share and a long tail captures slivers.

  • Words: "the," "of," "and" dominate; rare words form an enormous tail.
  • Cities: a handful of metropolises, then countless towns.
  • Firms: a few giant corporations, a sea of small businesses.
  • Web traffic: a few mega-sites, a vast long tail of small ones.
  • Income: the Pareto distribution, Zipf's economic cousin.

This recurrence across unrelated systems is why Zipf's law is often called "universal," though that label is contested. The shared trait is a process where a few items grow huge while most stay small, frequently through proportionate growth: everything grows at roughly the same percentage rate, so the gaps between ranks widen multiplicatively. Economist Xavier Gabaix gave the canonical proportionate-growth explanation for cities in 1999. The rank-frequency curve has a sister probability distribution, P(n) ∝ 1 / n^β, with β = 1 + 1/α, putting β near 2. If this kind of ranking fascinates you, the related idea of how value scales with connections appears in the network value calculator.

The competing "why" theories

No one fully agrees on the cause. Three families of explanation compete.

Least-effort communication

Zipf's own theory framed language as a tug-of-war. A speaker wants a small vocabulary of reusable words to minimize their own effort; a listener wants a large, precise vocabulary to minimize ambiguity. The 2003 Ferrer-i-Cancho and Sole model formalized this tension and showed that Zipf's law emerges precisely at the balance point between the two pressures. Experiments confirm that when both pressures (accuracy and efficiency) act in conflict, language users converge on Zipf-like lexicons; remove one pressure and the effect disappears.

Random letter combination

The deflating alternative is that no linguistic mechanism is needed at all. If you generate "words" by hitting random keys, with spaces appearing at random, the resulting gibberish still produces a Zipf-like frequency curve. This random-text result, associated with Benoit Mandelbrot and later researchers, suggests the law might be a statistical artifact of any token-and-separator system rather than evidence of communicative optimization. Mandelbrot also refined the formula itself, adding a constant to fit the high-frequency end better: f(r) ≈ C / (r + b)^α.

Sample-space reduction

A more recent mechanistic account, published in the Journal of the Royal Society Interface in 2015, derives Zipf's law from "sample-space reducing" processes. As a sentence unfolds, each word constrains the plausible next words, shrinking the space of options as you go. This nested narrowing alone reproduces the 1/rank scaling, needing no assumptions about preferential attachment or self-organized criticality, only the empirically measurable nestedness of word transitions.

The honest caveat: where it breaks down

Zipf's law is an approximation, not a hard rule, and it frays at both ends. Among the very most frequent function words, the fit often bulges above the straight line. In the long tail of rare words, the curve typically drops faster than 1/rank predicts. By Britannica's account the simple law largely breaks down beyond about rank 1,000, though that figure depends heavily on corpus size; analyses of Wikipedia have found Zipf-like behavior holding for the top 10,000 words, while studies of city sizes show clear deviations for the smallest and largest cities. The exact breakdown point is not fixed; it shifts with how much text you have and even how you define a "word" or a "city."

That caveat matters more than it sounds. A near-straight log-log line is also produced by a lognormal distribution, so a good visual fit does not by itself prove an underlying power law. The safest reading is that Zipf's law is a robust descriptive regularity with real explanatory puzzles behind it, not a settled law of nature. To watch the curve emerge from your own writing, paste any block of text into the Zipf's law calculator and compare the measured slope against the ideal negative-one line.

Frequently Asked Questions

Because Zipf's law says frequency is roughly inversely proportional to rank, so the rank-2 word appears about 1/2 as often as rank-1, rank-3 about 1/3 as often, and so on. The relationship is approximate, not exact, and works best for the more common words in a large text.

No. It is a strong empirical regularity, not a derived law. Several competing theories try to explain it, including least-effort communication, random text generation, and sample-space reduction, and none is universally accepted. It also fails at the extremes, so it is best treated as a useful approximation.

Yes, approximately. City populations, firm sizes, website traffic, and income distributions all show similar 1/rank patterns on a log-log plot with a slope near negative one. The fit is imperfect, especially for the very largest and smallest items, and depends on how the units are defined.

There is no fixed cutoff. Britannica notes the simple form largely breaks down past about rank 1,000, but larger corpora like Wikipedia can show Zipf-like behavior into the top 10,000 words. The breakdown point shifts with corpus size and how you define a word.