The Square Root Law of Lexical Evolution
For decades, historical linguistics assumed that basic vocabulary replaced itself at an approximately uniform rate — a glottochronologic constant of roughly 14% per millennium, proposed by Morris Swadesh in the 1950s. Data gathered across 87 Indo-European languages by Mark Pagel's team at the University of Reading demonstrates that replacement rates vary by more than a factor of 34 between words in the same language.
The determining factor is spoken frequency. Word frequency follows a Zipfian power law: a few dozen words constitute half of all daily speech, while the remainder of the dictionary comprises rare occurrences. The replacement rate of a word (λ) scales inversely with the square root of its usage frequency (f):
If a word is spoken 100 times more frequently than another, its half-life is 10 times longer. If spoken 10,000 times more frequently, its half-life increases by 100 times.
This relationship explains why a 21st-century time traveler listening to a conversation 6,000 years ago on the Pontic steppe would immediately recognize the numbers: *dwóh₁ became two in English, doi in Romanian, deux in French, dva in Russian, and dva in Sanskrit. Numbers and pronouns are spoken every few minutes; human neural networks reinforce them in each generation, preventing lexical drift.
By contrast, the adjective for "dirty" sits in the long frequency tail. Without the stabilizing pressure of continuous repetition, speech communities replaced the term repeatedly: English borrowed dirty from Norse drit ("dung"), French took sale from Frankish, Spanish derived sucio from Latin sucidus ("sweaty"), and Romanian borrowed murdar from Turkish. No modern language preserves a common shared root for this concept.
Extinction of Irregular Forms: The Fate of 177 Verbs
The exact same mathematical law governs grammatical morphology. In 2007, Harvard researchers (Lieberman, Michel, Jackson, Tang, and Nowak) analyzed the fate of 177 irregular verbs recorded in 9th-century Old English (the era of Beowulf).
As the language passed through Middle English (Chaucer) into Modern English, irregular verbs behaved like a biological population under selective pressure. Of the 177 original irregular verbs, 79 permanently regularized (44.6%), adopting the standard -ed past tense suffix.
| Verb (EN / RO) | Frequency / 1M | Old English (800) | Middle English (1300) | Modern Form | Half-Life |
|---|---|---|---|---|---|
| be (a fi) | 39,175 | wesan / bēon | been / was | irregular (was/been) | 38,800 years |
| have (a avea) | 12,450 | habban / hæfde | haven / hadde | irregular (had) | 28,400 years |
| do (a face) | 4,380 | dōn / dyde | doon / dide | irregular (did/done) | 17,200 years |
| say (a spune) | 3,120 | secgan / sæġde | sayen / seide | irregular (said) | 14,500 years |
| help (a ajuta) | 310 | healp / holpen | halp / holpen | regularized (helped) | 2,800 years |
| climb (a urca) | 55 | clamb / clumben | clomb | regularized (climbed) | 1,200 years |
| bake (a coace) | 34 | bacan / bōc | boke | regularized (baked) | 950 years |
| chew (a mesteca) | 18 | cēowan / cēaw | chaw | regularized (chewed) | 700 years |
The regularization rate followed the identical power law: rare verbs quickly collapsed into the regular rule, while the top 10 most frequent verbs (be, have, do, say, go, get, make, see, know, take) have an estimated resistance of 10,000 to 38,800 years.
In Romance languages like Romanian, the mechanism operates identically: all surviving irregular verbs are high-frequency Latin inheritances. Every newly coined or borrowed verb over the last millennium has entered productive regular classes. Zero newly formed verbs have acquired irregular inflection.
Irregular Verb Stability Calculator
Enter a verb's usage frequency (occurrences per million words) to compute its estimated half-life against regularization into the standard paradigm: