# Which language is closest to Romanian? Three measures, three answers

Romanian, the Romance language of Romania and Moldova, was compared with the world's languages using open data on basic words, grammar and speech sounds. The nearest languages differ from one measure to the next, and any language can be looked up.

Marius Comper · 4 October 2026 · data: ASJP, WALS, Grambank, PHOIBLE

## The maps of nearness

On the maps a dot is a language, placed where it is spoken. The larger and stronger the dot, the nearer the language is to Romanian on that measure. The empty circle marks Romanian. Sounds have two maps, one for each description of Romanian.

### Words

1. Megleno-Romanian: 0.44
2. Italian: 0.60
3. Sardinian (Logudorese): 0.61
4. Aromanian: 0.61
5. Neapolitan and southern Italian: 0.62
6. Emilian: 0.66
7. Sardinian (Campidanese): 0.66
8. Catalan: 0.67
9. Piemontese: 0.67
10. Friulian: 0.67

### Grammar

1. Italian: 41 of 51
2. Albanian: 45 of 56
3. Portuguese: 40 of 50
4. Greek: 48 of 65
5. Spanish: 47 of 64
6. Ukrainian: 29 of 40
7. Bulgarian: 41 of 57
8. Russian: 46 of 65
9. Gulf Arabic: 33 of 47
10. Catalan: 28 of 40

### Sounds: with 7 vowels and 20 consonants

1. Damana (Malayo): 1–11
2. Bangolan: 1–22
3. Nupe: 3–11
4. Karipúna French Creole: 2–17
5. Mazanderani: 1–56
6. Sayula Popoluca: 8–32
7. Ndamba: 4–37
8. Nigerian Pidgin: 3–54
9. Ishkashimi: 8–46
10. Kirundi: 8–52

### Sounds: with the softened consonants of Romanian “lupi” (wolves) and “pomi” (trees) counted separately

1. Russian: 1–5
2. Upper Sorbian: 3–7
3. Veps: 1–7
4. Skolt Saami: 1–10
5. Lithuanian: 1–12
6. Votic: 6–10
7. Erzya: 4–11
8. Lower Sorbian: 6–9
9. Bulgarian: 1–62
10. Ter Saami: 10–21

Three open databases describe the world's languages from three sides: basic words, grammar and sounds. Each was used the same way here: Romanian compared in turn with every language in the database.

Relatives, neighbours and languages that merely have the same sounds are three different things. Basic words are inherited from the parent language and change slowly. Grammar is inherited too, but over centuries of living side by side it moves towards the neighbours'. A language has few sounds, a few dozen, and two languages can end up with almost the same sounds without ever having met.

## One language beside Romanian

Type the name of a language. The four rulers start on the left, at Romanian: the further left a language's mark, the nearer it is on that measure. Under the rulers is the evidence: the 40 words side by side, the grammar features on which the two languages agree or differ, the sounds they share and those found in only one of the two.

Each rank comes with a range. The calculation can be done in several equally defensible ways (which word list for Romanian, which description of its sounds, how many shared features are required); the range gives the best and the worst rank the language takes across all of them. For sounds the median rank is given too, the middle one when the variants are put in order. There are 5,734 languages in all, but only 201 are ranked on all three measures.

## Words: the relatives come first

The ASJP project (Automated Similarity Judgment Program) has collected 11,540 lists from around the world, each built on the same 40 meanings, the most stable ones in a classic list of 100 (the Swadesh list): I, you, we, one, two, water, fire, sun, name. Each word is transcribed with a small set of phonetic symbols. The distance between two languages is worked out word by word: how many sounds must be changed, added or removed to get from one form to the other, relative to length. The result is then divided by the average distance between words of different meanings in the same two lists. Zero means identical forms for every meaning compared. Around 1 (the score can go slightly above it), words with the same meaning are no more alike than words picked at random.

The nearest list is the Megleno-Romanian one (0.44), in all 72 variants of the calculation. Then come Italian (0.60), Logudorese Sardinian (0.61), Aromanian (0.61, but with only 32 words of the 40 in common) and Neapolitan (0.62). In the baseline calculation the first 17 places are all taken by Romance languages (the last of them by a Portuguese-based creole of India). Spanish is in place 29 (0.73), French in 33 (0.78).

Megleno-Romanian is spoken in Greece and North Macedonia, Aromanian in those two countries, in Albania and in their neighbours. Romanian linguistics traditionally counts them as historical dialects of Romanian, together with the Romanian of Romania itself and with Istro-Romanian in Croatia; the international catalogues Glottolog and ISO 639-3 list them as separate languages. Whatever they are called, they are Romanian's closest relatives, and the first language outside that group is Italian. ASJP has no separate list for Istro-Romanian, nor for the speech of the Republic of Moldova, which Glottolog enters as a dialect of Romanian.

In the baseline calculation Latin comes only in place 18 (0.69); depending on the variant it is between places 10 and 25, but always after Italian and Logudorese Sardinian. The measure compares today's forms sound by sound and says nothing about descent. Romanian “foc” and Italian “fuoco”, fire, continue Latin “focus”, which meant hearth; for fire, ASJP's Latin list has “ignis”.

Beyond the Romance languages the scores quickly approach 1. Bulgarian is in place 53 (0.87), Albanian in 245 (0.93), Mandarin Chinese in 2,843 (0.99), Hungarian in 3,846 (1.00). More than half of the 5,457 languages score above 0.99, and only 127 fall below 0.90. There, where nearly every language crowds close to 1, a language's rank says very little. Nor does the ranking measure borrowed vocabulary: Bulgarian and Hungarian, from which Romanian has taken many words, are compared on 40 stable meanings only, without the words marked as loans.

ASJP also holds lists for constructed languages, left out of the ranking here. Interlingua, built from the shared vocabulary of the major European languages, most of it Romance, would be second (0.59); Esperanto would be in place 29 (0.72).

## Grammar: a relative and a neighbour score the same

WALS, the World Atlas of Language Structures, describes each language's grammar through features with a handful of possible values: where the definite article goes, how many cases a noun has, in what order subject, verb and object come. Romanian has 82 features described; 65 concern grammar, the rest sounds and vocabulary. The measure used here is simple: of the features described for both languages, how many have the same value.

Italian has the same value as Romanian on 41 of 51 features, Albanian on 45 of 56, Portuguese on 40 of 50: four in five, for all three. Then come Greek (48 of 65), Spanish (47 of 64), Ukrainian (29 of 40) and Bulgarian (41 of 57). French, a Romance language, has 38 of 64, fewer than Russian (46 of 65) or English (45 of 65). At the other end, Turkish: 21 of 58.

With 50 to 60 features, small differences cannot separate languages: the margins of error of the first ten overlap. What holds across the nine variants of the calculation: Albanian is in the first three in all of them, Italian is in the first five in eight (in the ninth it has too few shared features to be ranked), and French never climbs above place 14.

The comparison is cleaner on the same features for every language. There are 36 described for all eight languages in the table below. On those, Albanian agrees with Romanian on 32, Italian on 30, Portuguese on 29, Spanish on 28, Bulgarian on 24, Greek on 23 and French on 21.

Where Romanian parts from Italian, Albanian is on Romanian's side. On eight features Romanian and Italian have different values and Albanian is described too; on five of them Albanian has Romanian's value, and on one Italian's. The five: the definite article attached to the end of the word (Romanian “lupul”, Albanian “ujku”, the wolf), cases marked by endings, ordinal numerals all formed from the cardinal ones except “first”, the way commands and exhortations are formed, and subject–verb order. Bulgarian and Greek do not show the same thing: in the same comparison Bulgarian sides with Romanian three times and with Italian six, Greek four and five times. One of Bulgarian's six is the definite article itself: WALS counts the Bulgarian one as a separate word, although it too comes after the noun (“вълкът”, the wolf).

The result is not a test of the Balkan sprachbund, the name linguists give to the grammatical likenesses between Romanian, Albanian, Bulgarian, Macedonian and Greek. Of its classic traits, WALS has the article placed after the noun; the infinitive replaced by a subjunctive clause (Romanian says “vreau să plec”, word for word “I want that I leave”, for “I want to leave”), the future formed with “want” and the object repeated by a pronoun have no features of their own. The article, the infinitive and the future are shown, with examples, on the page on the Balkan sprachbund. Italian is not far from Albanian either: the two have the same value on 42 of 58 features described for both.

### What Grambank says, without Romanian

Grambank, the largest comparative grammar database according to its authors (2,467 language varieties, 195 features), has no record for Romanian in version 1.0.3. It has none for Bulgarian or for Spanish either. It has Aromanian, with 165 features filled in, so the test can be repeated for Romanian's closest relative that has data.

As Grambank codes them, Aromanian agrees most often with Italian and with Portuguese (151 of 164 features each), then with Occitan, Galician, Lombard and Corsican. Tosk Albanian, the basis of standard Albanian, comes right after them (142 of 160), ahead of Sardinian, Catalan (142 of 164) and French (136 of 165).

Here too, where Aromanian parts from Italian, Albanian is on Aromanian's side: nine times against four. Macedonian gives 9 to 3, Greek 7 to 6. Among the features: the article placed after the noun, cases, then mood and tense marked by invariable particles. Two of these codings were challenged in review: Grambank records Italian as having no auxiliary verb for tense, although “ho visto” has one, and the second concerns the pronoun attached after the verb in Aromanian. Without them the scores are 7 to 4 for Albanian, 7 to 3 for Macedonian and 5 to 6 for Greek.

## Sounds: two descriptions of Romanian, two lists

PHOIBLE collects sound inventories. An inventory is the list of a language's sounds that can tell two words apart (its phonemes, sometimes with their variants), as a linguist has established it. There are 3,020 inventories for 2,186 languages. What is measured here is how many sounds two inventories have in common, out of all the sounds they have between them; one variant of the calculation measures instead how close the sounds of one inventory are to those of the other. It does not measure how often each sound occurs in speech, nor the melody of a sentence, nor how the language sounds to someone who does not know it.

For Romanian, PHOIBLE has three inventories, and they do not agree. Two follow the usual description: 7 vowels, 20 consonants and the semivowels of “iarnă” (winter), “ziuă” (day), “seară” (evening) and “soare” (sun). The third, compiled from the linguist Ioana Chitoran's 2002 book on the sound system of Romanian, counts separately the softened (palatalised) consonants at the end of words such as “lupi” (wolves), “pomi” (trees), “rupi” (you break): the p of “lupi” is entered as a different sound from the p of “lup” (wolf). The inventory thus reaches 51 segments, against 31 and 32.

With the usual description, the nearest inventories belong to distant languages: Damana in Colombia, Bangolan in Cameroon, Nupe in Nigeria, a French-based creole of Brazil (Karipúna) and Mazanderani in Iran. The first three have no connection to Romanian; Mazanderani is an Iranian language, so a very distant relative. Like Romanian, they have a medium-sized inventory made mostly of sounds found all over the world. Bangolan has all 29 sounds present in the three descriptions of Romanian: seven vowels, 20 consonants and the semivowels of “iarnă” and “ziuă”.

When the softened consonants are counted separately, the list moves to northern and eastern Europe: Russian, Sorbian, Veps, Skolt Saami, Lithuanian, Erzya, Bulgarian. All have whole series of softened consonants, as Russian does in “брать” (brat', to take) against “брат” (brat, brother).

In none of the 72 variants of the calculation does Italian climb above place 19, and Mandarin Chinese never climbs above place 986, of 2,085. Albanian has 28 of the 29 sounds common to the descriptions of Romanian, Italian 25 (it lacks “ă”, “î”, “h” and the “j” of “joc”), Bulgarian 24, Mandarin 17.

Of those 29 sounds, the rarest in PHOIBLE's inventories are the “j” of “joc” (game), the sound of the s in “measure” (ʒ, present in 15.9% of inventories), “î” (ɨ, 16.5%) and “ă” (ə, 22.4%). Then come “ț” (ts, 26.8%) and “v” (27.0%). The most widespread sound missing from all three descriptions of Romanian is the “ng” of English “sing” (ŋ, in 62.8% of inventories). Among the next are the “ñ” of Spanish (ɲ, 42.7%) and the glottal stop, the catch in the middle of “uh-oh” (ʔ, 37.5%).

How Romanian sounds to a stranger's ear is a different question. In an online game in which people guessed a language from a 20-second recording, with 15 million guesses analysed in 2017 by Hedvig Skirgård, Seán Roberts and Lars Yencken, on the map of the players' confusions Romanian sits next to Greek, Albanian and the Slavic languages.

## Why the three answers do not agree

Each measure follows something else. The list of 40 words is made of the most stable meanings, and the first 17 places in the baseline calculation are all taken by Romance languages, Romanian's relatives.

Grammar is inherited, but it also moves towards the languages around it. Romanian is spoken far from the other Romance languages: Slavic languages and Hungarian lie between it and Italian, while Bulgarian, Albanian and Greek are nearer on the map. On the chart the Romance languages are near on both measures; Albanian, Greek and Bulgarian are near on grammar and far on words.

Sounds are few. A language has a few dozen phonemes, and the most widespread (p, t, k, m, n, a, i, u) turn up almost everywhere. Two medium-sized inventories can overlap by more than three quarters without the languages having anything to do with each other, as with Romanian and Damana.

That is why the question in the title has no single answer. “Closest” means something only together with the measure: on words, on grammar or on sounds.

## How it was measured and which choices were made

### Language identity

A language here is an entry at “language” level in the Glottolog 5.3 catalogue. Lists and inventories collected for dialects are assigned to their language. Standard Albanian rests on the Tosk dialect. ASJP's list for standard Albanian, attached to the group “Albanian”, and WALS's record “Albanian” (which has the ISO code sqi, no Glottolog code, and mentions both the Gheg and the Tosk dialect) were placed with Tosk Albanian, because PHOIBLE and Grambank keep their Albanian data under that code.

### Words

ASJP version 21: 11,540 lists. The baseline calculation follows the project's rules: the 40-word list, at most two synonyms for a meaning, words marked as loans set aside, the Levenshtein distance adjusted for word length and divided by the average distance between words of different meanings (LDND). Romanian has three lists, which differ from one another by 0.12 to 0.22; the baseline is ROMANIAN_3. A language is ranked if it shares at least 28 words with the Romanian list: 5,457 languages. Left out: 405 lists with no Glottolog code, 22 lists of constructed languages, four lists attached to groups of languages and 620 languages with too few words.

The variants: the Romanian list (3), compound symbols taken as one unit or split (2), loans removed or kept (2), synonyms (the first two, all, or the closest pair: 3), the median list or the nearest when a language has several (2). 72 in all. The word table shows the median list; when that is the list of a local variety and ASJP also has a list named plainly after the language, the plain one is shown, with its own distance. For Mandarin Chinese, which has 184 lists, the table shows the Beijing list.

### Grammar

WALS, edition 2020.4: 65 features of Romanian, after removing the 15 on phonology and the two on vocabulary; 28 of them describe word order. The baseline ranking requires at least 40 shared features: 243 languages, of the 1,900 languages that share at least one. The variants: all features, without the word-order ones (37), or in blocks, where features that follow from one another count once (51 blocks); with the threshold at 30, 40 or 50 shared features.

Grambank 1.0.3 was searched for Romanian's Glottolog code and for the whole Eastern Romance branch: the only record is Aromanian's. Aromanian was compared with the languages that share at least 120 features with it.

### Sounds

PHOIBLE 2.0.1. Before comparison the notations were brought to one form: the marks for dental, apical, laminal, lowered, raised, centralised, advanced and retracted were removed; tapped r and trilled r were counted as one r; the semivowels written i̯ and u̯ were read as j and w. Tones, diphthongs and the semivowels of “seară” and “soare” are left out. Length, nasalisation, palatalisation, aspiration and voicing remain differences. This is a deliberate simplification: without it the same Romanian t would be three different sounds in the three descriptions.

The variants: the measure (overlap on the simplified notation, overlap on PHOIBLE's own notation, or a distance on PHOIBLE's 37 phonetic features: 3), diphthongs left out or kept (2), marginal phonemes kept or removed (2), the description of Romanian (3), the median inventory or the nearest when a language has several (2). 72 in all. In the sound-by-sound comparison a Romanian sound counts as shared if at least one inventory of the chosen language has it as an ordinary phoneme.

### What is missing

No measure here says how easily two speakers understand each other, and none takes account of the whole vocabulary, of spelling or of actual pronunciation. The databases are uneven: WALS has few features for many languages, and PHOIBLE mixes descriptions made to different conventions. That is why every result is given together with the number of words or features it rests on.

## Sources

- [Wichmann, Holman, Brown, Dryer & Ran (eds.), The ASJP Database, version 21 (2025)](https://doi.org/10.5281/zenodo.16736409) The word lists; CLDF release 21.1, licence CC BY 4.0.
- [Dryer & Haspelmath (eds.), The World Atlas of Language Structures Online (2013), v2020.4](https://doi.org/10.5281/zenodo.13950591) The grammar features; CC BY 4.0. Romanian's record: wals.info/languoid/lect/wals_code_rom.
- [Grambank v1.0.3 (2023)](https://doi.org/10.5281/zenodo.7844558) The grammar features of Aromanian and of the languages compared with it; CC BY 4.0.
- [Skirgård, Haynie, Blasi, Hammarström, Collins et al., “Grambank reveals the importance of genealogical constraints on linguistic diversity…”, Science Advances 9(16), 2023](https://doi.org/10.1126/sciadv.adg6175) The paper describing Grambank.
- [Moran & McCloy (eds.), PHOIBLE 2.0 (2019), Max Planck Institute for the Science of Human History](https://phoible.org/) The sound inventories; CLDF release 2.0.1, licence CC BY-SA 3.0. Romanian's inventories: phoible.org/languages/roma1327.
- [Hammarström, Forkel, Haspelmath & Bank, Glottolog 5.3 (2026)](https://doi.org/10.5281/zenodo.18840935) Language identity, classification and coordinates; CC BY 4.0.
- [Natural Earth, 1:110m land](https://www.naturalearthdata.com/) The land outline on the maps; public domain.
- [Wichmann, Holman, Bakker & Brown, “Evaluating linguistic distance measures”, Physica A 389(17), 2010, 3632–3639](https://doi.org/10.1016/j.physa.2010.05.011) The LDND distance used for words.
- [ASJP, software and instructions](https://asjp.clld.org/help) The rules followed: at most two synonyms, loans excluded.
- [Chitoran, The Phonology of Romanian: A Constraint-Based Approach, Mouton de Gruyter, 2002](https://doi.org/10.1515/9783110889185) The source of the inventory with palatalised consonants (EA 2443 in PHOIBLE).
- [Agard, “Structural Sketch of Rumanian”, Language 34(3, part 2), 1958, 7–127](https://doi.org/10.2307/522282) One source of the other two inventories of Romanian (SPA 165, UPSID 527).
- [Lee & Zee, “Standard Chinese (Beijing)”, Journal of the International Phonetic Association 33(1), 2003, 109–112](https://doi.org/10.1017/S0025100303001208) The source of two of the four Standard Chinese inventories in PHOIBLE.
- [Skirgård, Roberts & Yencken, “Why are some languages confused for others? Investigating data from the Great Language Game”, PLOS ONE 12(4), 2017](https://doi.org/10.1371/journal.pone.0165934) Which languages Romanian is confused with by ear.
- [Anderson, Tresoldi, Greenhill, Forkel, Gray & List, “Variation in phoneme inventories: quantifying the problem and improving comparability”, Journal of Language Evolution 8(2), 2023, 149–168](https://doi.org/10.1093/jole/lzad011) Why inventories of the same language differ from one database to the next.
- [Littell, Mortensen, Lin, Kairis, Turner & Levin, “URIEL and lang2vec”, EACL 2017](https://doi.org/10.18653/v1/E17-2002) A library of distances between languages built from the same sources.
