# The Half-Life of Words

> Like radioactive isotopes, words in human languages have a measurable mathematical half-life: words spoken 100 times more frequently evolve and get replaced 10 times more slowly.

- **Canonical URL**: https://mariuscomper.uk/jumatatea-de-viata/en/
- **Romanian Edition**: https://mariuscomper.uk/jumatatea-de-viata/
- **Author**: Marius Comper (2026)

---

## Key Metrics and Evolutionary Constants

- **10×**: Slowdown in replacement rate for every 100-fold increase in usage frequency ($f^{-0.5}$ law).
- **26,000 years**: Calculated half-life for the pronoun "who" (*kʷi-), conserved across nearly all branches.
- **750 years**: Half-life for the adjective "dirty", replaced independently in Romanian, French, and English.
- **44.6%**: Percentage of the 177 Old English irregular verbs regularized to "-ed" over the last 1,200 years (79 verbs).

---

## The Square Root Law of Lexical Evolution

For decades, historical linguistics assumed that basic vocabulary replaced itself at an approximately uniform rate — a glottochronologic constant of roughly 14% per millennium, proposed by Morris Swadesh in the 1950s. Data gathered across 87 Indo-European languages by Mark Pagel's team at the University of Reading demonstrates that replacement rates vary by more than a factor of 34 between words in the same language.

The determining factor is spoken frequency. Word frequency follows a Zipfian power law: a few dozen words constitute half of all daily speech, while the remainder of the dictionary comprises rare occurrences. The replacement rate of a word ($\lambda$) scales inversely with the square root of its usage frequency ($f$):

$$\lambda \propto \frac{1}{\sqrt{f}} \quad \Longrightarrow \quad t_{1/2} \propto \sqrt{f}$$

If a word is spoken 100 times more frequently than another, its half-life is 10 times longer. If spoken 10,000 times more frequently, its half-life increases by 100 times.

This relationship explains why a 21st-century time traveler listening to a conversation 6,000 years ago on the Pontic steppe would immediately recognize the numbers: *dwóh₁* became **two** in English, **doi** in Romanian, **deux** in French, **dva** in Russian, and **dva** in Sanskrit. Numbers and pronouns are spoken every few minutes; human neural networks reinforce them in each generation, preventing lexical drift.

By contrast, the adjective for "dirty" sits in the long frequency tail. Without the stabilizing pressure of continuous repetition, speech communities replaced the term repeatedly: English borrowed *dirty* from Norse *drit* ("dung"), French took *sale* from Frankish, Spanish derived *sucio* from Latin *sucidus* ("sweaty"), and Romanian borrowed *murdar* from Turkish. No modern language preserves a common shared root for this concept.

---

## Extinction of Irregular Forms: The Fate of 177 Verbs

The exact same mathematical law governs grammatical morphology. In 2007, Harvard researchers (Lieberman, Michel, Jackson, Tang, and Nowak) analyzed the fate of 177 irregular verbs recorded in 9th-century Old English (the era of *Beowulf*).

As the language passed through Middle English (Chaucer) into Modern English, irregular verbs behaved like a biological population under selective pressure. Of the 177 original irregular verbs, 79 permanently regularized (44.6%), adopting the standard *-ed* past tense suffix.

| Verb (EN / RO) | Frequency / 1M | Old English (800) | Middle English (1300) | Modern Form | Half-Life |
| :--- | :--- | :--- | :--- | :--- | :--- |
| **be** (a fi) | 39,175 | wesan / bēon | been / was | irregular (was/been) | 38,800 years |
| **have** (a avea) | 12,450 | habban / hæfde | haven / hadde | irregular (had) | 28,400 years |
| **do** (a face) | 4,380 | dōn / dyde | doon / dide | irregular (did/done) | 17,200 years |
| **say** (a spune) | 3,120 | secgan / sæġde | sayen / seide | irregular (said) | 14,500 years |
| **help** (a ajuta) | 310 | healp / holpen | halp / holpen | regularized (helped) | 2,800 years |
| **climb** (a urca) | 55 | clamb / clumben | clomb | regularized (climbed) | 1,200 years |
| **bake** (a coace) | 34 | bacan / bōc | boke | regularized (baked) | 950 years |
| **chew** (a mesteca) | 18 | cēowan / cēaw | chaw | regularized (chewed) | 700 years |

The regularization rate followed the identical power law: rare verbs quickly collapsed into the regular rule, while the top 10 most frequent verbs (*be, have, do, say, go, get, make, see, know, take*) have an estimated resistance of 10,000 to 38,800 years.

In Romance languages like Romanian, the mechanism operates identically: all surviving irregular verbs are high-frequency Latin inheritances. Every newly coined or borrowed verb over the last millennium has entered productive regular classes. Zero newly formed verbs have acquired irregular inflection.

---

## Method Note & References

This project synthesizes lexical evolution measurements published by Mark Pagel, Quentin D. Atkinson, and Andrew Meade in *Nature* (vol. 449, 2007) with the historical analysis of irregular verb dynamics by Erez Lieberman, Jean-Baptiste Michel, Joe Jackson, Tina Tang, and Martin A. Nowak in *Nature* (vol. 449, 2007).

Replacement rates $\lambda$ represent the average number of lexical substitution events per 1,000 years of linguistic history. Half-life is calculated using the exponential decay equation: $t_{1/2} = \frac{\ln(2)}{\lambda} \times 1,000\text{ years}$. Reconstructed Proto-Indo-European roots and cognates follow standard Indo-European etymological lexicons.
