Language Kinetics & Glottochronology

The Half-Life of Words

Like radioactive isotopes, words in human languages have a measurable mathematical half-life: words spoken 100 times more frequently evolve and get replaced 10 times more slowly. While core pronouns and numbers endure unchanged for over 20,000 years, rare terms decay every few centuries.

10×
Slowdown in replacement rate for every 100-fold increase in usage frequency (f−0.5 law).
26,000
Years is the calculated half-life for the pronoun "who" (*kʷi-), conserved across almost all branches.
750
Years is the half-life for the adjective "dirty", replaced independently in Romanian, French, and English.
44.6%
Of the 177 Old English irregular verbs regularized to "-ed" over the last 1,200 years.

Lexical Decay Chamber

Move the timeline slider from the emergence of Proto-Indo-European roots into the future. Each word computes its survival probability P(t) = 2−Δt / t½ based on spoken frequency.

2026 AD
Loading data...
4000 BC (Proto-Indo-European) 1500 BC (Sanskrit / Mycenaean) 100 BC (Classical Latin) 800 AD (Old English) 2026 (Today) 4000 AD (Future)

The Square Root Law of Lexical Evolution

For decades, historical linguistics assumed that basic vocabulary replaced itself at an approximately uniform rate — a glottochronologic constant of roughly 14% per millennium, proposed by Morris Swadesh in the 1950s. Data gathered across 87 Indo-European languages by Mark Pagel's team at the University of Reading demonstrates that replacement rates vary by more than a factor of 34 between words in the same language.

The determining factor is spoken frequency. Word frequency follows a Zipfian power law: a few dozen words constitute half of all daily speech, while the remainder of the dictionary comprises rare occurrences. The replacement rate of a word (λ) scales inversely with the square root of its usage frequency (f):

Kinetic Vocabulary Equation (Pagel et al., 2007)
λ ∝ 1 / √f   ⟹   t½ ∝ √f

If a word is spoken 100 times more frequently than another, its half-life is 10 times longer. If spoken 10,000 times more frequently, its half-life increases by 100 times.

This relationship explains why a 21st-century time traveler listening to a conversation 6,000 years ago on the Pontic steppe would immediately recognize the numbers: *dwóh₁ became two in English, doi in Romanian, deux in French, dva in Russian, and dva in Sanskrit. Numbers and pronouns are spoken every few minutes; human neural networks reinforce them in each generation, preventing lexical drift.

By contrast, the adjective for "dirty" sits in the long frequency tail. Without the stabilizing pressure of continuous repetition, speech communities replaced the term repeatedly: English borrowed dirty from Norse drit ("dung"), French took sale from Frankish, Spanish derived sucio from Latin sucidus ("sweaty"), and Romanian borrowed murdar from Turkish. No modern language preserves a common shared root for this concept.

Extinction of Irregular Forms: The Fate of 177 Verbs

The exact same mathematical law governs grammatical morphology. In 2007, Harvard researchers (Lieberman, Michel, Jackson, Tang, and Nowak) analyzed the fate of 177 irregular verbs recorded in 9th-century Old English (the era of Beowulf).

As the language passed through Middle English (Chaucer) into Modern English, irregular verbs behaved like a biological population under selective pressure. Of the 177 original irregular verbs, 79 permanently regularized (44.6%), adopting the standard -ed past tense suffix.

Verb (EN / RO) Frequency / 1M Old English (800) Middle English (1300) Modern Form Half-Life
be (a fi) 39,175 wesan / bēon been / was irregular (was/been) 38,800 years
have (a avea) 12,450 habban / hæfde haven / hadde irregular (had) 28,400 years
do (a face) 4,380 dōn / dyde doon / dide irregular (did/done) 17,200 years
say (a spune) 3,120 secgan / sæġde sayen / seide irregular (said) 14,500 years
help (a ajuta) 310 healp / holpen halp / holpen regularized (helped) 2,800 years
climb (a urca) 55 clamb / clumben clomb regularized (climbed) 1,200 years
bake (a coace) 34 bacan / bōc boke regularized (baked) 950 years
chew (a mesteca) 18 cēowan / cēaw chaw regularized (chewed) 700 years

The regularization rate followed the identical power law: rare verbs quickly collapsed into the regular rule, while the top 10 most frequent verbs (be, have, do, say, go, get, make, see, know, take) have an estimated resistance of 10,000 to 38,800 years.

In Romance languages like Romanian, the mechanism operates identically: all surviving irregular verbs are high-frequency Latin inheritances. Every newly coined or borrowed verb over the last millennium has entered productive regular classes. Zero newly formed verbs have acquired irregular inflection.

Irregular Verb Stability Calculator

Enter a verb's usage frequency (occurrences per million words) to compute its estimated half-life against regularization into the standard paradigm:

occurrences / 1M
Estimated half-life against regularization: ≈ 6,196 years

Method Note & References

This project synthesizes lexical evolution measurements published by Mark Pagel, Quentin D. Atkinson, and Andrew Meade in Nature (vol. 449, 2007) with the historical analysis of irregular verb dynamics by Erez Lieberman, Jean-Baptiste Michel, Joe Jackson, Tina Tang, and Martin A. Nowak in Nature (vol. 449, 2007).

Replacement rates λ represent the average number of lexical substitution events per 1,000 years of linguistic history. Half-life is calculated using the exponential decay equation: t½ = ln(2) / λ × 1,000 years. Reconstructed Proto-Indo-European roots and cognates follow standard Indo-European etymological lexicons.