Tea Took Two Roads
opening the database
← mariuscomper.uk
RO · EN

Lexical atlas · 107 languages · 1,016 words

Tea took two roads

The map behind this text shows the word for tea in one hundred and seven languages of Eurasia, each one written where its language is spoken. The orange forms came overland, by caravan. The blue ones came by sea, in Dutch ships. Both start from the same Chinese character, and the line between them is one of the oldest trade borders on the continent.

Read on ↓

I · Tea

The word the continent shares most widely

Linguists at the University of Tübingen assembled the same core vocabulary in one hundred and seven languages of Eurasia, under the name NorthEuraLex: Greenlandic, Tamil, Basque, Ainu, Chechen, Sakha, Romanian. One thousand and sixteen concepts, from water and mother to getting tired and the rainbow, all written in phonetic script, which is to say in a form that can be compared. The result is 121,611 words laid side by side.

Of all 1,016 concepts, the one the continent agrees on most widely is tea. One hundred and six languages have a word for it, and seventy-two of them, spread across twenty unrelated language families, use forms close enough to fall into a single group.

The explanation is a trade map. Chinese writes the plant with one character, 茶, but reads it differently by province. Whoever bought tea overland, through Central Asia, heard the northern form and passed it on: chá in Persian, çay in Turkish, чай in Russian, ceai in Romanian. Whoever bought it in port, from Fujian merchants, heard the southern form and carried it by sea: thee in Dutch, then tea, thé, Tee, te. Seventy-four languages in the database are on the land road and twenty-eight on the water road.

Polish and Lithuanian give a third answer, herbata and arbata, from the Latin of apothecaries, herba thea, the tea herb: there the plant arrived through the pharmacy rather than through trade.

The explorer

Type a word

The field searches all 1,016 concepts. The map redraws on every choice, the colours mark groups of forms that resemble one another, and words that no longer fit are left as dots. Move a finger or the cursor across the map to see which language each one is.

The figure beside each group is the number of languages in it.

II · The calendar

After tea come the months

The ranking of concepts the continent shares most widely is led by tea, and directly below it, one after another, sit the months of the Roman calendar. March holds the same form in sixty-six languages, April and May in sixty, August and February in fifty-eight, October in fifty-four. The names of Latin gods and emperors, carried as far as Kamchatka by an administrative system for dividing the year.

    The concepts with the largest group of similar forms, out of 1,016. The figure on the right is how many languages fall in the largest group, out of those that have a word for the concept. Tap any row to send it to the explorer.

    The rest of the ranking is made of the same material. Master and newspaper, the cross, and a little further down salt, table and soup. These are things that travelled with merchants, with missionaries and with print, and that reached new languages carrying their names along.

    Nothing near the top of the ranking is inherited. What the continent shares is not the vocabulary it received from its ancestors. It is the vocabulary it bought from its neighbours.

    III · The cuckoo

    The one exception is a bird

    Somewhere among the top places sits a concept nobody bought. One hundred and five languages have a name for the cuckoo, and seventy-three of them, from thirteen families, call it almost the same thing.

    Finnish says käki, Hungarian kakukk, German Kuckuck, French coucou, Welsh cwcw, Turkish guguk, Russian кукушка, Greenlandic kukkooq, Itelmen in Kamchatka ӄэӄуӄ, Basque kuku, Romanian cuc. Basque is related to nobody, Itelmen is eleven thousand kilometres away from it, and both arrived at the same syllable.

    Tap cuckoo to see its map in the explorer.

    This is neither inheritance nor borrowing. It is a bird that calls out its own name twice every spring, loudly enough to be heard across the northern hemisphere, and each language writes it down with the sounds it has to hand. When people have nowhere to take a word from, they make one, and they make it out of the same material.

    IV · What is not shared

    On the rainbow nobody agrees

    The ranking also has a tail, and the tail is far longer than the head. At the midpoint of all 1,016 concepts, the largest group of similar forms covers under a tenth of the languages. Five hundred and twenty-one concepts fall below that line.

      The concepts that break into the most groups of forms. The figure on the right is how many distinct groups the words of the hundred-odd languages that have one fall into.

      The list is nearly uniform. To recover, to hurry, to shine, to deceive, sometimes, to fall ill, to sparkle, to get tired. These are things that happen inside somebody's body or head, and that you cannot point at in order to ask a neighbour what they call it.

      An object travels with its name, because whoever buys tea buys the word too. Tiredness is not for sale. Each language built it out of whatever was in the house, usually from an image borrowed from somewhere else, and the images chosen do not match.

      The one thing on the list you can point at is also the most striking case. One hundred and sixty-two concepts in the database have a word in all one hundred and seven languages, and the most shattered among them is the rainbow: eighty-four groups of forms, well ahead of the next. A rainbow is visible from three counties at once and nobody bought one from anybody, so each language described it on its own, and the descriptions do not resemble each other.

      The continent shares what it bought and shares almost nothing of what it feels.

      V · Romanian

      Who resembles Romanian most

      The same database allows a smaller question. Out of 1,016 concepts, in how many does Romanian use a word that resembles another particular language's. The order at the top is the expected one, and the first six places are all Romance.

        How many of the 1,016 concepts take a word in Romanian resembling the one in the listed language. Blue marks the Slavic languages.

        Seventh place is the interesting part. Bulgarian comes ahead of French, with Croatian close behind it. Romanian spent a thousand years in a Slavic neighbourhood and took from it a layer of vocabulary that covers precisely the things you do with your hands and around a household.

        Forty-nine concepts where Romanian goes with the Slavs and with no Romance language

        These are tools, illnesses, marsh animals and states of mind: the shovel, the sledge, the broom, the hook, the wound, the remedy, the squirrel, the swan, the pike, the goose, the spirit, the voice, the blame, cheerful, lazy, weak. Nothing to do with the yard, the neighbours or misfortune stayed Latin.

        One hundred and eighteen concepts where Romanian goes with the Romance languages and with no Slavic one

        Here are the body, counting, farm animals and writing. What Latin held and what it gave up shows plainly: to measure, to calculate, to write, to buy, to sell, the cow, the bull, the fish, the ant, the hay, the iron, the name, the people, the number, the arm.

        One hundred and seventy-four concepts where Romanian resembles nobody in the database

        The last list calls for caution, because part of it is the limit of the method rather than the history of the language. Romanian says apă where Latin said aqua, and the two really are the same word, except that the sound became p along the way, and a measure that compares consonants cannot cross a distance like that. The same happens with limbă against lingua. In other cases Romanian genuinely does stand alone, because it kept a word the rest of the family replaced, or because it took one from the substrate that preceded Latin.

        VI · The islands

        When the neighbour beats the relative

        If you ask, for each language in the database, which of the other hundred and six is closest to it, nearly all of them answer with a relative. Twelve answer with a language from another family.

          The twelve languages whose closest lexical neighbour comes from another family, with the percentage of the 1,016 concepts in which the two use similar words. The figure on the left is how many relatives the language has in the database.

          For nine of them the answer was settled in advance, since they have no relative in the database at all: Basque, Georgian, Korean, Japanese, Ainu, Nivkh, Ket, Burushaski and Chinese are the sole representatives of their families here. The question only becomes meaningful for the other three.

          Hungarian is the clearest case. It has twenty-five Uralic relatives in the database, among them Khanty and Mansi from beyond the Urals, and the language it shares the most everyday words with is Slovak. Lezgian, from Daghestan, leaves behind its five Nakh-Daghestanian relatives and answers Azerbaijani. Itelmen, from Kamchatka, leaves behind Chukchi and answers Russian, the language that arrived there last.

          All three are small languages caught between large neighbours, and this vocabulary is the everyday kind, where words change most easily. On a word list chosen specifically for how hard it is to borrow, Hungarian would answer Mansi.

          Method

          The database is NorthEuraLex 0.9, assembled by Johannes Dellert and colleagues at the University of Tübingen and published under a CC BY-SA 4.0 licence. It holds 121,611 forms across 1,016 concepts and 107 languages, each also written in the International Phonetic Alphabet. The language coordinates are those in the database and mark an approximate centre of the area where each is spoken, not a boundary.

          The grouping of forms is computed here rather than taken from the source. Each word is reduced to a string of sound classes, in the manner proposed by Aharon Dolgopolsky: p, b and f fall into one class, as do t with d and k with g, because between distant languages these turn into one another. Vowels carry little weight, glides adjacent to a vowel are absorbed into it, and the first consonant has to match, being the most stable part of a word. Two forms join the same group when the distance between their strings falls below a threshold, and groups merge by average linkage, so that a chain of small resemblances cannot gather together words with nothing in common.

          What comes out of this is resemblance, not genealogy. Resemblance most often comes from inheritance or from borrowing, but it sometimes comes from coincidence, and the method cannot tell them apart. In the other direction it misses real relatives wherever the sounds have drifted too far, as with apă against aqua. The figures on this page read as a measure of similarity in form, and the maps show where similar forms cluster rather than where a language family ends.

          Concepts with at least sixty forms were kept, which is all 1,016, and the rankings use only concepts covered in at least eighty languages. The Romanian labels in the explorer are the words Romanian itself supplies in the database, in its dictionary forms.

          Sources

          • Johannes Dellert et al., NorthEuraLex: a wide-coverage lexical database of Northern Eurasia, Language Resources and Evaluation 54, 2020. northeuralex.org, licensed CC BY-SA 4.0.
          • Aharon Dolgopolsky, Gipoteza drevnejšego rodstva jazykovyx semej Severnoj Evrazii, 1964, for the sound classes.
          • Land outline: Natural Earth, public domain.