# Tea, thé, чай, 茶: who says it like you?

https://mariuscomper.uk/harta-cuvintelor/en/

Pick a word and see how 107 languages say it, from Iceland to Kamchatka. The ones that say it almost like you light up first.

Your language

Loading the map…

How close each word sounds to yours

sounds like yours · nothing in common

Tea in the languages of the map, grouped by words that sound alike:

Finnish *tee*, North Karelian *čäijy*, Olonets Karelian *čuaju*, Veps *čai*, Estonian *tee*, Livonian *tēj*, Lule Sami *tedja*, Northern Sami *deadju*, Inari Sami *čee*, Skolt Sami *čee*, Kildin Sami *чайй*, Hill Mari *чӓй*, Meadow Mari *чай*, Moksha *чай*, Erzya *чай*, Udmurt *чай*, Komi-Permyak *чай*, Komi-Zyrian *чай*, Hungarian *tea*, Northern Khanty *шай*, Northern Mansi *ся̄й*, Northern Selkup *чой*, Tundra Nenets *сяй*, Forest Enets *чай*, Nganasan *чаи*, Bengali *চা*, Hindi *चाय*, Northern Pashto *چای*, Western Farsi *چای*, Northern Kurdish *çay*, Ossetian *цай*, Armenian *թեյ*, Standard Albanian *çaj*, Bulgarian *чай*, Croatian *čaj*, Slovene *čáj*, Slovak *čaj*, Czech *čaj*, Ukrainian *чай*, Belarusian *чай*, Russian *чай*, Latvian *tēja*, Icelandic *te*, Norwegian (Bokmål) *te*, Swedish *te*, Danish *te*, German *Tee*, Dutch *thee*, English *tea*, Irish *tae*, Welsh *te*, Breton *te*, French *thé*, Catalan *te*, Spanish *té*, Portuguese *chá*, Italian *tè*, Romanian *ceai*, Turkish *çay*, North Azerbaijani *çay*, Northern Uzbek *choy*, Kazakh *шай*, Bashkir *сәй*, Tatar *чәй*, Sakha *чэй*, Chuvash *чей*, Khalkha Mongolian *цай*, Buryat *сай*, Kalmyk *цә*, Evenki *чай*, Manchu *cai*, Nanai *чаи*, Northern Yukaghir *чай*, Itelmen *чай*, Chukchi *чай*, Nivkh *чай*, Hokkaido Ainu *cha*, Korean *차*, Japanese *茶*, Kalaallisut *tii*, Kannada *ಚಹ*, Malayalam *ചായ*, Burushaski *ćái*, Georgian *ჩაი*, Basque *te*, Adyghe *щай*, Avar *чай*, Tsez *чай*, Lak *чяй*, Lezgian *чай*, Dargwa *чай*, Chechen *чай*, Standard Arabic *شاي*, Modern Hebrew *תה*, Mandarin Chinese *茶*

Modern Greek *τσάι*, Latin *thea*

Polish *herbata*, Lithuanian *arbata*

Aleut *chaayu-x́*, Central Siberian Yupik *ӄаюӄ*

Tamil *தெனீர்*, Telugu *తేనీరు*

One colour means one group: words that resemble the whole group, and between different language families the likeness has also passed a chance test. They may be inherited, borrowed, imitated from a sound or, sometimes, a coincidence. Orange is your group. The strip under the map places each language by sound alone: on the left, words almost the same as yours; on the right, words with nothing in common. Tap a dot to see the word and how it is pronounced.

Try:

1 · Tea

## Tea reached Europe by two roads

Russians say *чай*, Turks *çay*, Persians *چای*, and in Hindi it is *चाय*. Max Vasmer’s etymological dictionary traces *чай* and the Turkic and Mongolic forms to northern Chinese *čhā*. Romanian took “ceai” from Russian, the DEX says.

The English say *tea*, the French *thé*, the Germans *Tee*. Dutch traders, Europe’s main importers of tea, dealt through Xiamen in Fujian province, where Min Nan is spoken and tea is *te*. From them the word spread to large parts of Europe, writes the linguist Östen Dahl.

Portuguese is the exception: the Portuguese brought tea before the Dutch, in the 16th century, through Macao, and say *chá*, from Cantonese.

The map puts all these words in one group, 95 languages from 20 families: the largest group of any of the 1,016 concepts. That grouping makes sense: both forms come from Chinese, *cha* from Mandarin and most Chinese languages, *te* from Min Nan. It still misses a few, from Greek *τσάι* to Polish *herbata*. The roads also show in the order the languages light up: starting from “tea”, mostly the words like *thé* and *Tee* come first, then the ones like *чай*.

## Around the world, the same two families of words

For the World Atlas of Language Structures, Dahl looked up the word for tea in 230 languages on every continent. 110 have a form from *cha*, 84 one from *te*, and 36 a word of another origin.

Languages of the countries in eastern Europe and Asia that got their tea overland, rather than from the Dutch, tend to say *chai*, Dahl writes. The Romanian academy’s dictionary makes the same split between kinds of tea: “Russian” tea was brought from China overland, through Russia; “English” tea by sea, through England.

2 · Mother and father

## Mother sounds alike in languages that are not related

In 79 languages out of 107, the first consonant of the word for mother is an *m*, an *n* or a soft *n*, like the *ny* of Hungarian *anya*: *mamă*, *мать*, *anya*, *ana*, *amma*, *ama*. They come from 16 language families.

In 1959 the anthropologist George Peter Murdock counted the words for mother and father in hundreds of languages and found this tendency: syllables like *ma* and *na* for mother, *pa* and *ta* for father, even in languages with no link between them. The linguist Roman Jakobson tied the tendency to children’s first syllables, which adults then gave a meaning.

## In Georgian, *mama* means father

For father, in 58 languages out of 107 the first consonant is *p*, *b*, *t* or *d*. Georgian does the opposite: father is *მამა* and mother is *დედა*. The map now starts from Georgian.

Murdock’s tendency stays a tendency: children’s syllables are alike everywhere, but languages gave them different meanings.

3 · The cuckoo

## The cuckoo says its own name

*Cuckoo*, *kakukk*, *guguk*, *kukkooq* in Greenland, *kuku* in Basque. On the map, 75 languages from 13 families have a name for the cuckoo like the English one.

It is not just a shared inheritance. Etymological dictionaries call the word imitative: Vasmer sets Turkish, Mordvin, Tatar and Kazakh forms beside Russian *кукушка*, from languages unrelated to Russian. Many of these names imitate the bird’s call.

4 · March

## Rome’s calendar reached Kamchatka

The month names of the Roman calendar are Latin: *Martius* was the month of Mars. Russian took the name from Latin through Byzantine Greek, Vasmer says. On the map, 68 languages from 14 families say something similar, as far as Itelmen in Kamchatka. In 20 other languages, from Bulgarian and Tatar to Chechen and Kalmyk, the word is spelled exactly as in Russian: *март*.

Some languages kept names of their own: Finnish *maaliskuu*, Czech *březen*, Ukrainian *березень*, Belarusian *сакавік*, Lithuanian *kovas*. The database’s authors included the months precisely because they are easily borrowed, as a test for methods that look for loanwords.

5 · The moon

## One word, two things

In Romanian, “lună” is both the moon in the sky and the month in the calendar. So it is in 57 other languages on the map, from 13 families. They are the ones lit up now.

Romanian “limbă” is both the tongue and the language, as in 56 other languages. In 53 languages, tree and wood share a name; in 45 languages, hand and arm. Linguists call this colexification. The CLICS database, which gathers more than 3,000 languages, finds the moon and the month together in 324 of them.

6 · Blue

## Where blue ends

With colours, the question becomes where the word stops. In 9 languages on the map, green and blue have the same name, according to NorthEuraLex: Yakut *күөх*, Ossetian *цъӕх*, Ainu *shiunin*, Adyghe, Kurdish and a few more.

The World Color Survey showed 330 coloured chips to speakers of 110 languages, about 24 per language. Draw the line where, for you, green becomes blue.

green · blue

A screen cannot show the survey’s Munsell chips exactly, so your line is an approximation. The full story is at [Where blue ends](https://albastru.mariuscomper.uk/en/).

## More than half the survey’s languages have one word

Paul Kay and Luisa Maffi’s WALS map has 120 languages, from the survey and related studies. 68 have one word for green and blue, 30 have separate words, and 15 have a word that also covers black. Most are unwritten languages, very different from the 107 on the map above.

7 · The rainbow

## The rainbow has dozens of names

Of the 218 concepts that all 107 languages have, the rainbow is among the most divided: 86 groups of words. Many languages describe it: German *Regenbogen*, the rain’s bow; French *arc-en-ciel*, the sky’s arc; Hindi *इंद्र-धनुष*, Indra’s bow.

The map puts Romanian “curcubeu” in the same group as Italian *arcobaleno*: both have r, c and b in the same order. Arcobaleno is *arco* and *baleno*, bow and flash. For “curcubeu”, Romanian dictionaries disagree: the 1998 DEX calls its origin unknown, the 2009 edition only compares it with Latin *curvus*, bent. The likeness looks like a coincidence of the kind described in the method below: between related languages, words often share a pattern of sounds.

Next time you ask for tea, you use the word the Dutch spread; Russians, Turks and Persians use the one that came overland. In Romanian and 57 other languages, the word for the moon also names the month. And when two words look alike, like “curcubeu” and *arcobaleno*, it is worth asking whether it is only chance.

[Pick another word](https://mariuscomper.uk/harta-cuvintelor/en/#harta)

## How the map decides what sounds alike

The words come from NorthEuraLex 0.9, a database made at the University of Tübingen: 1,016 concepts in 107 languages, each word also written in the phonetic alphabet. The transcriptions were mostly generated automatically from spelling, and the authors warn that the lists contain errors.

The map groups words by sound. Between different language families it shows only likenesses that are hard to put down to chance. Which of them are also related is for linguists to establish, from each word’s history. The map takes three steps.

### 1. Compare the sounds

Each word is reduced to sound classes, so that *p* and *b* count as one sound and *f* as very close to them, then aligned with the others. The method, called SCA, is the linguist Johann-Mattis List’s and runs in the LingPy software.

### 2. Check whether the likeness could be chance

With 106 languages to compare, almost any word sounds like something. We checked: we replaced a language’s word with another word of the same language that means something else. On sound alone, this false word found a lookalike in an unrelated language family in 53 of 100 tries.

So the map works in two steps. First it groups the words within each language family by sound alone: relatives share many words, and this is the comparison List and his colleagues tested against experts’ decisions.

Then it joins groups from different families, where a match comes mostly from borrowing, from imitating a sound, from children’s syllables, as in *mama*, or from chance. For every pair of languages, the likeness is compared with 4,000 pairs of words picked at random from the same two languages, with different meanings; for this test, sounds are reduced to even simpler classes. One pair proves little: a short word like “tea” can resemble another by pure chance. The same match repeated across many language families is hard to put down to chance. The map asks that the likeness, taken over all the family pairs involved, be rarer than 2 in 1,000; the probabilities are combined with Fisher’s method.

With this rule, the false word still finds a lookalike in another family in 3 of 100 tries. In its own family it finds one in 23 of 100, because relatives have words with the same pattern of sounds; there, the right benchmark is the comparison with experts below.

### 3. Group

Words join a group only if they resemble the whole group on average, not just one member. The colour on the map shows the group. The strip under the map shows something else: only how close each word sounds to yours, without the chance test. That is why a hollow dot can sit near you on the strip: it sounds alike, but the likeness may be a coincidence.

### How often it is right

For the Indo-European languages there is a benchmark: IE-CoR, a database in which historical linguists decided, for 160 languages, which words share an origin. We compared the map’s groups with their decisions on 4,129 words from 33 languages present in both databases and 156 concepts; the false-word trials used the same concepts. When the map puts two words in one group, the experts consider them related in 84% of cases. Of the pairs the experts link, the map finds 50%.

Many of the pairs it misses are relatives that no longer look alike: English *fish* and Romanian *pește*, or *four* and *patru*, go back to the same Indo-European words, according to the IE-CoR experts. The map measures likeness today, not history. It groups inherited words and borrowed ones alike, such as Romanian “ceai” and Russian *чай*.

### Whom English sounds like

Counting the concepts where two languages have words in the same group:

- Language · Concepts in the same group · Of

- Dutch · 358 · 1,016
- German · 333 · 1,016
- Swedish · 331 · 1,016
- Norwegian (Bokmål) · 323 · 1,015
- Danish · 301 · 1,015
- Icelandic · 252 · 1,016
- French · 192 · 1,016
- Italian · 181 · 1,016

Family counts treat each isolate, such as Basque or Ainu, as a family of its own. “One form for two meanings”, for the moon or the tongue, means an identical NorthEuraLex transcription for both concepts. NorthEuraLex is one of the sources of CLICS, so the two counts are not independent.

## Sources

- [Dellert et al., “NorthEuraLex: a wide-coverage lexical database of Northern Eurasia”, Language Resources and Evaluation, 2020](https://doi.org/10.1007/s10579-019-09480-6). Data: [northeuralex.org](http://www.northeuralex.org/), version 0.9, CC BY-SA 4.0.

- [Johann-Mattis List, “LexStat: Automatic detection of cognates in multilingual wordlists”, 2012](https://aclanthology.org/W12-0216/), and [List, Greenhill and Gray, “The potential of automatic word comparison for historical linguistics”, PLOS ONE, 2017](https://doi.org/10.1371/journal.pone.0170046).

- [IE-CoR, Indo-European Cognate Relationships](https://iecor.clld.org/); [Anderson et al., Scientific Data, 2025](https://doi.org/10.1038/s41597-025-05445-3).

- [Östen Dahl, “Tea”, World Atlas of Language Structures](https://wals.info/chapter/138); [Paul Kay and Luisa Maffi, “Green and Blue”](https://wals.info/chapter/134).

- [Rzymski et al., “The Database of Cross-Linguistic Colexifications”, Scientific Data, 2020](https://doi.org/10.1038/s41597-019-0341-x).

- [George Peter Murdock, “Cross-language parallels in parental kin terms”, 1959, HRAF summary](https://hraf.yale.edu/ehc/documents/745).

- Max Vasmer, Etymological Dictionary of the Russian Language, entries [чай](https://gufo.me/dict/vasmer/чай), [кукушка](https://gufo.me/dict/vasmer/кукушка) and [март](https://gufo.me/dict/vasmer/март).

- Romanian dictionaries (DEX, MDA2, Ciorănescu) via [dexonline](https://dexonline.ro/): ceai, cuc, curcubeu, lună, limbă.

- [Where blue ends](https://albastru.mariuscomper.uk/en/): the World Color Survey boundaries, recomputed from the original data.

[mariuscomper.uk · Marius Comper](https://mariuscomper.uk/en/) · First version: August 2026; rebuilt in September 2026
