mariuscomper.uk ROEN中文

Romanian and the world's languages

Which language is closest to Romanian? Three measures, three answers

Romanian, the Romance language of Romania and Moldova, was compared with the world's languages using open data on basic words, grammar and speech sounds. The nearest languages differ from one measure to the next, and any language can be looked up.

  • WordsMegleno-Romanian, then Italian
  • GrammarItalian, Albanian and Portuguese, all at about four in five
  • SoundsRussian and other languages of northern and eastern Europe, or two small languages of Colombia and Cameroon, depending on which description of Romanian’s sounds is used

The maps of nearness

On the maps a dot is a language, placed where it is spoken. The larger and stronger the dot, the nearer the language is to Romanian on that measure. The empty circle marks Romanian. Sounds have two maps, one for each description of Romanian.

Words

40 basic words, compared sound by sound, in the 5,457 languages that have at least 28 of them.

World map: the languages nearest to Romanian on words cluster around the Mediterranean.World map: the languages nearest to Romanian on words cluster around the Mediterranean.
  1. Megleno-Romanian0.44
  2. Italian0.60
  3. Sardinian (Logudorese)0.61
  4. Aromanian0.61
  5. Neapolitan and southern Italian0.62

The first 17 places are all taken by Romance languages.

Grammar

65 features of grammar (where the article goes, how many cases a noun has, the order of words), in the 243 languages that have at least 40 of them described.

World map: the languages nearest to Romanian on grammar are in Europe.World map: the languages nearest to Romanian on grammar are in Europe.
  1. Italian41 of 51
  2. Albanian45 of 56
  3. Portuguese40 of 50
  4. Greek48 of 65
  5. Spanish47 of 64

The first three score practically the same: a relative, a neighbour, a relative.

Sounds

How many sounds each language has in common with Romanian, in 2,085 languages. The list depends on the description of Romanian.

With 7 vowels and 20 consonants

World map: the languages whose sets of speech sounds are nearest to Romanian's on this measure are scattered across Africa, the Americas and South-East Asia.World map: the languages whose sets of speech sounds are nearest to Romanian's on this measure are scattered across Africa, the Americas and South-East Asia.
  1. Damana (Malayo) Colombia1–11
  2. Bangolan Cameroon1–22
  3. Nupe Nigeria3–11
  4. Karipúna French Creole Brazil2–17
  5. Mazanderani Iran1–56

With the softened consonants of Romanian “lupi” (wolves) and “pomi” (trees) counted separately

World map: with the softened consonants counted separately, the nearest languages are in northern and eastern Europe.World map: with the softened consonants counted separately, the nearest languages are in northern and eastern Europe.
  1. Russian1–5
  2. Upper Sorbian Germany3–7
  3. Veps Russia1–7
  4. Skolt Saami Finland and Russia1–10
  5. Lithuanian1–12

Italian is outside the first 18 in every variant.

Three open databases describe the world's languages from three sides: basic words, grammar and sounds. Each was used the same way here: Romanian compared in turn with every language in the database.

Relatives, neighbours and languages that merely have the same sounds are three different things. Basic words are inherited from the parent language and change slowly. Grammar is inherited too, but over centuries of living side by side it moves towards the neighbours'. A language has few sounds, a few dozen, and two languages can end up with almost the same sounds without ever having met.

One language beside Romanian

Type the name of a language. The four rulers start on the left, at Romanian: the further left a language's mark, the nearer it is on that measure. Under the rulers is the evidence: the 40 words side by side, the grammar features on which the two languages agree or differ, the sounds they share and those found in only one of the two.

Italian

Indo-European family · 986 km from Romania · Glottolog: Italian [ital1282]

Words rank 2 of 5,457 · between 2 and 3 across all variants · distance 0.60 · words compared: 40 of 40

RomanianItalianMegleno-RomanianFrenchBulgarianHungarian

Grammar rank 1 of 243 · between 1 and 5 across all variants · features with the same value: 41 of 51

RomanianItalianEnglishHungarianTurkish

Sounds, the usual description median rank 128 of 2,085 (between 19 and 279, depending on the variant) · sounds shared with Romanian: 25 of 29

RomanianItalianAlbanianRussianMandarin Chinese

Sounds, with the soft consonants median rank 158 of 2,085 (between 39 and 376, depending on the variant)

RomanianItalianRussianBulgarianMandarin Chinese

← near Romanianfar →

The 40 words

ASJP list: ITALIAN_2 · distance 0.60 · words compared: 40 of 40

MeaningRomanianItalianDistance
Ijewio1.00
you (one person)tutu0.00
wenojnoi0.33
oneunuuno0.33
twodojdue0.67
personom, persoanəuomo, ɛssere umano0.73
fishpeʃtepeʃʃe0.20
dogkəjnekane0.40
lousepədukepidokkjo0.63
treekopak, arborealbero0.75
leaffrunzəfoʎʎa0.83
skinpjelepelle0.40
bloodsəndʒesaŋgue0.67
boneososso0.50
hornkornkorno0.20
earurekeorekkjo0.57
eyeokʲokkjo0.80
nosenasnazo0.50
toothdintedente0.20
tonguelimbəliŋgua0.67
kneedʒenunkʲdʒinokkjo0.75
handmənəmano0.50
breastsənseno0.50
liverfikatfegato0.50
to drinkbeabere0.50
to seevedeavedere0.33
to hearawzʲudire, sentire1.00
to diemurʲmorire0.83
to comevenʲvenire0.67
sunsoaresole0.40
starsteastella0.33
waterapəakkwa0.80
stonepjatrəsasso, pietra0.67
firefokfuoko0.40
pathkəraresentiero0.88
mountainmuntemonte, montaɲɲa0.41
nightnoaptenotte0.33
fullplinpieno0.60
newnownuovo0.60
namenumenome0.25
Grammar, feature by feature

Different values (10)

  • Definite ArticlesRomanian: Definite affixItalian: Definite word distinct from demonstrative
  • Politeness Distinctions in PronounsRomanian: Multiple politeness distinctionsItalian: Binary politeness distinction
  • Indefinite PronounsRomanian: Interrogative-basedItalian: Generic-noun-based
  • Number of CasesRomanian: 2 casesItalian: No morphological case-marking
  • Position of Case AffixesRomanian: Case suffixesItalian: No case affixes or adpositional clitics
  • Ordinal NumeralsRomanian: First, two-th, three-thItalian: First, second, three-th
  • Imperative-Hortative SystemsRomanian: Maximal systemItalian: Neither type of system
  • Order of Subject and VerbRomanian: SVItalian: No dominant order
  • Order of Demonstrative and NounRomanian: MixedItalian: Demonstrative-Noun
  • Negative Indefinite Pronouns and Predicate NegationRomanian: Predicate negation also presentItalian: Mixed behaviour

Same value (41)

  • Prefixing vs. Suffixing in Inflectional MorphologyStrongly suffixing
  • Coding of Nominal PluralityPlural suffix
  • Occurrence of Nominal PluralityAll nouns, always obligatory
  • Indefinite ArticlesIndefinite word same as 'one'
  • Intensifiers and Reflexive PronounsDifferentiated
  • Asymmetrical Case-MarkingAdditive-quantitatively asymmetrical
  • Comitatives and InstrumentalsIdentity
  • Position of Pronominal Possessive AffixesNo possessive affixes
  • Position of Tense-Aspect AffixesTense-aspect suffixes
  • The Morphological ImperativeSecond singular
  • The ProhibitiveSpecial imperative + normal negative
  • Situational PossibilityVerbal constructions
  • Epistemic PossibilityVerbal constructions
  • Overlap between Situational and Epistemic Modal MarkingOverlap for both possibility and necessity
  • Order of Subject, Object and VerbSVO
  • Order of Object and VerbVO
  • Order of Adposition and Noun PhrasePrepositions
  • Order of Genitive and NounNoun-Genitive
  • Order of Adjective and NounNoun-Adjective
  • Order of Numeral and NounNumeral-Noun
  • Order of Relative Clause and NounNoun-Relative clause
  • Postnominal relative clausesNoun-Relative clause (NRel) dominant
  • Order of Degree Word and AdjectiveDegree word-Adjective
  • Position of Polar Question ParticlesNo question particle
  • Order of Adverbial Subordinator and ClauseInitial subordinator word
  • Relationship between the Order of Object and Verb and the Order of Adposition and Noun PhraseVO and Prepositions
  • Relationship between the Order of Object and Verb and the Order of Relative Clause and NounVO and NRel
  • Relationship between the Order of Object and Verb and the Order of Adjective and NounVO and NAdj
  • Negative MorphemesNegative particle
  • Polar QuestionsInterrogative intonation only
  • Order of Negative Morpheme and VerbNegV
  • Preverbal Negative MorphemesNegV
  • Postverbal Negative MorphemesNone
  • Minor morphological means of signaling negationNone
  • Position of Negative Word With Respect to Subject, Object, and VerbSNegVO
  • Position of negative words relative to beginning and end of clause and with respect to adjacency to verbImmed preverbal
  • The Position of Negative Morphemes in SVO LanguagesSNegVO
  • NegSVO OrderNo NegSVO
  • SNegVO OrderWord&NoDoubleNeg
  • SVNegO OrderNo SVNegO
  • SVONeg OrderNo SVONeg

Grambank, Aromanian against this language: features with the same value, 151 of 164.

The sounds

Shared sounds (25)

abddʒefijklmnoprsttstʃuvwzɡʃ

Only in Romanian (4)

həɨʒ

Only in Italian (5)

dzɔɛɲʎ

inventories in PHOIBLE: 3, sizes from 30 to 70

Each rank comes with a range. The calculation can be done in several equally defensible ways (which word list for Romanian, which description of its sounds, how many shared features are required); the range gives the best and the worst rank the language takes across all of them. For sounds the median rank is given too, the middle one when the variants are put in order. There are 5,734 languages in all, but only 201 are ranked on all three measures.

Words: the relatives come first

The ASJP project (Automated Similarity Judgment Program) has collected 11,540 lists from around the world, each built on the same 40 meanings, the most stable ones in a classic list of 100 (the Swadesh list): I, you, we, one, two, water, fire, sun, name. Each word is transcribed with a small set of phonetic symbols. The distance between two languages is worked out word by word: how many sounds must be changed, added or removed to get from one form to the other, relative to length. The result is then divided by the average distance between words of different meanings in the same two lists. Zero means identical forms for every meaning compared. Around 1 (the score can go slightly above it), words with the same meaning are no more alike than words picked at random.

The nearest list is the Megleno-Romanian one (0.44), in all 72 variants of the calculation. Then come Italian (0.60), Logudorese Sardinian (0.61), Aromanian (0.61, but with only 32 words of the 40 in common) and Neapolitan (0.62). In the baseline calculation the first 17 places are all taken by Romance languages (the last of them by a Portuguese-based creole of India). Spanish is in place 29 (0.73), French in 33 (0.78).

  1. Megleno-Romanian0.44
  2. Italian0.60
  3. Sardinian (Logudorese)0.61
  4. Aromanian0.61
  5. Neapolitan and southern Italian0.62
  6. Emilian0.66
  7. Sardinian (Campidanese)0.66
  8. Catalan0.67
  9. Piemontese0.67
  10. Friulian0.67
  11. Sicilian0.67
  12. Lombard0.67
  13. Sassarese0.68
  14. Galician0.68
  15. Corsican0.68
  16. Portuguese0.68
  17. Korlai Portuguese Creole0.69
  18. Latin0.69
  19. Romansh0.69
  20. Ligurian0.69
  21. Asturian-Leonese0.70
  22. Romagnol0.70
  23. Aragonese0.70
  24. Dalmatian0.70
  25. Ladin0.71
  26. Occitan0.71
  27. Daman and Diu Portuguese Creole0.71
  28. Mirandese0.72
  29. Spanish0.73
  30. Ladino0.76
The 30 nearest languages on words. The dot is the distance in the baseline calculation; the line spans the distances in all 72 variants.

Megleno-Romanian is spoken in Greece and North Macedonia, Aromanian in those two countries, in Albania and in their neighbours. Romanian linguistics traditionally counts them as historical dialects of Romanian, together with the Romanian of Romania itself and with Istro-Romanian in Croatia; the international catalogues Glottolog and ISO 639-3 list them as separate languages. Whatever they are called, they are Romanian's closest relatives, and the first language outside that group is Italian. ASJP has no separate list for Istro-Romanian, nor for the speech of the Republic of Moldova, which Glottolog enters as a dialect of Romanian.

In the baseline calculation Latin comes only in place 18 (0.69); depending on the variant it is between places 10 and 25, but always after Italian and Logudorese Sardinian. The measure compares today's forms sound by sound and says nothing about descent. Romanian “foc” and Italian “fuoco”, fire, continue Latin “focus”, which meant hearth; for fire, ASJP's Latin list has “ignis”.

Beyond the Romance languages the scores quickly approach 1. Bulgarian is in place 53 (0.87), Albanian in 245 (0.93), Mandarin Chinese in 2,843 (0.99), Hungarian in 3,846 (1.00). More than half of the 5,457 languages score above 0.99, and only 127 fall below 0.90. There, where nearly every language crowds close to 1, a language's rank says very little. Nor does the ranking measure borrowed vocabulary: Bulgarian and Hungarian, from which Romanian has taken many words, are compared on 40 stable meanings only, without the words marked as loans.

0.40.60.81.0Megleno-RomanianItalianFrenchBulgariandistance from Romanian
All 5,457 languages, by distance from Romanian. Nearly all of them gather around 1.

ASJP also holds lists for constructed languages, left out of the ranking here. Interlingua, built from the shared vocabulary of the major European languages, most of it Romance, would be second (0.59); Esperanto would be in place 29 (0.72).

The 40 words, side by side

ASJP writes sounds with 41 symbols, shown here in the International Phonetic Alphabet; Romanian “ă” and “î” share one symbol (ə). The lists are shown as ASJP has them, with their inconsistencies of transcription. The small number beside each word is its distance from the Romanian word: 0 for identical forms, 1 when every sound has to be changed. The Aromanian list is the one ASJP calls VLACH and has 32 words of the 40.

MeaningRomanianMegleno-RomanianAromanianItalianLatinFrench
Ijewjo, jew 0.33eu, siŋuru 0.83io 1.00ego 1.00ʒə, mwa 1.00
you (one person)tutu 0.00tini 0.75tu 0.00tu 0.00ti, twa 0.58
wenojnoj 0.00noi 0.33noi 0.33nos 0.33nu 0.67
oneunuun 0.33unu, una 0.17uno 0.33unus 0.25ɛ 1.00
twodojdoj 0.00dui, doua 0.58due 0.67duo 0.67de 0.67
personom, persoanəwom 0.60omu, birbatu 0.74uomo, ɛssere umano 0.73persona, homo 0.62om 0.44
fishpeʃtepeaʃti 0.33pesku 0.60peʃʃe 0.20piskis 0.83pwaso 0.80
dogkəjnekojni 0.40–kane 0.40kanis 0.80ʃjɛ 0.80
lousepədukepidukʎu 0.43–pidokkjo 0.63pedikulus 0.67pu 0.67
treekopak, arborearbur 0.67arbor, pomu 0.70albero 0.75arbor 0.58aχbχə 0.83
leaffrunzəfrunzə 0.00–foʎʎa 0.83folʲũ 0.83fɛj 0.83
skinpjelekoaʒə 1.00kiali 0.80pelle 0.40kutis 1.00po 0.80
bloodsəndʒesonzi 0.60sintsa 0.60saŋgue 0.67saŋgʷis 0.83so 0.80
boneoswos 0.33–osso 0.50os 0.00os 0.00
hornkornkorn 0.00–korno 0.20kornu 0.20koχn 0.25
earurekeureakʎə 0.43uriatsli 0.71orekkjo 0.57auris 0.80oχɛj 1.00
eyeokʲwokʎu 0.80okklu 0.80okkjo 0.80okulus 0.83ɛj 1.00
nosenasnas 0.00–nazo 0.50nasus 0.40ne 0.67
toothdintedinti 0.20dinti 0.20dente 0.20dens 0.60do 0.80
tonguelimbəlimbe 0.20limba 0.20liŋgua 0.67liŋgʷɛ 0.60log 0.80
kneedʒenunkʲzinkʎu 0.83tsinuklu 0.71dʒinokkjo 0.75genu 0.50ʒənu 0.67
handmənəmonə 0.25mina 0.50mano 0.50manus 0.60mɛ 0.75
breastsəntʃiept, sin 0.67keptu 1.00seno 0.50pektus, mama 1.00sɛ 0.67
liverfikatdrop 1.00–fegato 0.50jekur 0.80fwa 0.60
to drinkbeabeaw 0.25abea, beaun 0.33bere 0.50bibere 0.67bwaχ 0.50
to seevedeavet 0.60vedu 0.40vedere 0.33widere 0.67vwaχ 0.80
to hearawzʲut 1.00aseultu, andu 0.80udire, sentire 1.00audire 0.83otodχə 1.00
to diemurʲmor 0.67moru 0.75morire 0.83mori 0.75muχiχ 0.60
to comevenʲvin 0.67gini, jinu 1.00venire 0.67wenire 0.83vəniχ 0.80
sunsoaresoari 0.20suari 0.40sole 0.40sol 0.60solɛj 0.60
starsteastewə 0.40stiaua 0.50stella 0.33stela 0.20etwal 0.60
waterapəapu 0.33apa 0.33akkwa 0.80akʷa 0.67o 1.00
stonepjatrəropə 0.83gena, osu 1.00sasso, pietra 0.67lapis 0.83pjɛχ 0.67
firefokfok 0.00foku 0.25fuoko 0.40iŋnis 1.00fe 0.67
pathkəraredrum 0.83–sentiero 0.88vija 0.83sotje 0.83
mountainmuntemunti 0.20mundi, dʒana 0.60monte, montaɲɲa 0.41mons 0.60mo, motaɲ 0.80
nightnoaptenoapti 0.17napte, lopti 0.33notte 0.33noks 0.67nji 0.83
fullplinplin, amplin 0.17mplinu 0.33pieno 0.60plenus 0.50plɛ, χopli 0.55
newnownow 0.00nou 0.33nuovo 0.60nowus 0.40nɛf, nuvo 0.71
namenumenumi 0.25numa 0.25nome 0.25nomen 0.40no 0.75
Distance over the whole list0.440.610.600.690.78

Grammar: a relative and a neighbour score the same

WALS, the World Atlas of Language Structures, describes each language's grammar through features with a handful of possible values: where the definite article goes, how many cases a noun has, in what order subject, verb and object come. Romanian has 82 features described; 65 concern grammar, the rest sounds and vocabulary. The measure used here is simple: of the features described for both languages, how many have the same value.

Italian has the same value as Romanian on 41 of 51 features, Albanian on 45 of 56, Portuguese on 40 of 50: four in five, for all three. Then come Greek (48 of 65), Spanish (47 of 64), Ukrainian (29 of 40) and Bulgarian (41 of 57). French, a Romance language, has 38 of 64, fewer than Russian (46 of 65) or English (45 of 65). At the other end, Turkish: 21 of 58.

With 50 to 60 features, small differences cannot separate languages: the margins of error of the first ten overlap. What holds across the nine variants of the calculation: Albanian is in the first three in all of them, Italian is in the first five in eight (in the ninth it has too few shared features to be ranked), and French never climbs above place 14.

  1. Italian41 of 51
  2. Albanian45 of 56
  3. Portuguese40 of 50
  4. Greek48 of 65
  5. Spanish47 of 64
  6. Ukrainian29 of 40
  7. Bulgarian41 of 57
  8. Russian46 of 65
  9. Gulf Arabic the Persian Gulf33 of 47
  10. Catalan28 of 40
  11. Polish39 of 56
  12. Hebrew45 of 65
  13. English45 of 65
  14. Woleaian Micronesia28 of 41
  15. Estonian34 of 50
  16. Bari South Sudan29 of 44
  17. Muna Indonesia27 of 41
  18. Khana Nigeria26 of 40
  19. Icelandic35 of 54
  20. Serbian-Croatian30 of 47
The first 20 languages on grammar. The dot is the share of features with the same value as Romanian; the line is the margin of error (Wilson confidence interval, 95%), indicative only, because the features are not independent of one another.

The comparison is cleaner on the same features for every language. There are 36 described for all eight languages in the table below. On those, Albanian agrees with Romanian on 32, Italian on 30, Portuguese on 29, Spanish on 28, Bulgarian on 24, Greek on 23 and French on 21.

On the same 36 features

  1. Albanian32
  2. Italian30
  3. Portuguese29
  4. Spanish28
  5. Bulgarian24
  6. Greek23
  7. French21

Where Romanian parts from Italian, Albanian is on Romanian's side. On eight features Romanian and Italian have different values and Albanian is described too; on five of them Albanian has Romanian's value, and on one Italian's. The five: the definite article attached to the end of the word (Romanian “lupul”, Albanian “ujku”, the wolf), cases marked by endings, ordinal numerals all formed from the cardinal ones except “first”, the way commands and exhortations are formed, and subject–verb order. Bulgarian and Greek do not show the same thing: in the same comparison Bulgarian sides with Romanian three times and with Italian six, Greek four and five times. One of Bulgarian's six is the definite article itself: WALS counts the Bulgarian one as a separate word, although it too comes after the noun (“вълкът”, the wolf).

Where Romanian and Italian differ, whose side Albanian is on

FeatureRomanianItalianAlbanian
Definite ArticlesDefinite affixDefinite word distinct from demonstrativeDefinite affixas Romanian
Position of Case AffixesCase suffixesNo case affixes or adpositional cliticsCase suffixesas Romanian
Ordinal NumeralsFirst, two-th, three-thFirst, second, three-thFirst, two-th, three-thas Romanian
Imperative-Hortative SystemsMaximal systemNeither type of systemMaximal systemas Romanian
Order of Subject and VerbSVNo dominant orderSVas Romanian
Order of Demonstrative and NounMixedDemonstrative-NounDemonstrative-Nounas Italian
Politeness Distinctions in PronounsMultiple politeness distinctionsBinary politeness distinctionNo politeness distinctionanother value
Number of Cases2 casesNo morphological case-marking4 casesanother value

The result is not a test of the Balkan sprachbund, the name linguists give to the grammatical likenesses between Romanian, Albanian, Bulgarian, Macedonian and Greek. Of its classic traits, WALS has the article placed after the noun; the infinitive replaced by a subjunctive clause (Romanian says “vreau să plec”, word for word “I want that I leave”, for “I want to leave”), the future formed with “want” and the object repeated by a pronoun have no features of their own. The article, the infinitive and the future are shown, with examples, on the page on the Balkan sprachbund. Italian is not far from Albanian either: the two have the same value on 42 of 58 features described for both.

What Grambank says, without Romanian

Grambank, the largest comparative grammar database according to its authors (2,467 language varieties, 195 features), has no record for Romanian in version 1.0.3. It has none for Bulgarian or for Spanish either. It has Aromanian, with 165 features filled in, so the test can be repeated for Romanian's closest relative that has data.

As Grambank codes them, Aromanian agrees most often with Italian and with Portuguese (151 of 164 features each), then with Occitan, Galician, Lombard and Corsican. Tosk Albanian, the basis of standard Albanian, comes right after them (142 of 160), ahead of Sardinian, Catalan (142 of 164) and French (136 of 165).

  1. Italian151 of 164
  2. Portuguese151 of 164
  3. Occitan150 of 164
  4. Galician149 of 164
  5. Lombard148 of 164
  6. Corsican139 of 156
  7. Albanian142 of 160
  8. Sicilian141 of 159
  9. Romansh139 of 158
  10. Sardinian (Campidanese)137 of 158
  11. Catalan142 of 164
  12. Ukrainian124 of 145
  13. Macedonian134 of 161
  14. Latvian133 of 160
  15. Belarusian135 of 163
  16. French136 of 165
  17. Danish128 of 156
  18. Gheg Albanian130 of 159
  19. Greek134 of 165
  20. Icelandic129 of 159

Here too, where Aromanian parts from Italian, Albanian is on Aromanian's side: nine times against four. Macedonian gives 9 to 3, Greek 7 to 6. Among the features: the article placed after the noun, cases, then mood and tense marked by invariable particles. Two of these codings were challenged in review: Grambank records Italian as having no auxiliary verb for tense, although “ho visto” has one, and the second concerns the pronoun attached after the verb in Aromanian. Without them the scores are 7 to 4 for Albanian, 7 to 3 for Macedonian and 5 to 6 for Greek.

All 65 features, in eight languages

The WALS values for Romanian and for seven languages of comparison. Marked cells have the same value as Romanian; a dash means the feature is not described for that language. In the order formulas S is the subject, V the verb, O the object, Neg the negation, and square brackets show a negation attached to the verb. Feature 45A, on polite pronouns, has its own page: Why We Say “Dumneavoastră” (the polite “you” of Romanian).

Show the table
FeatureRomanianItalianFrenchSpanishPortugueseAlbanianBulgarianGreek
26A Prefixing vs. Suffixing in Inflectional MorphologyDoes the language add mostly beginnings or endings to words?Strongly suffixingStrongly suffixingStrongly suffixingStrongly suffixingStrongly suffixingStrongly suffixingStrongly suffixingStrongly suffixing
33A Coding of Nominal PluralityHow does the language show that a noun is plural?Plural suffixPlural suffixPlural suffixPlural suffixPlural suffixPlural suffixPlural suffixPlural suffix
34A Occurrence of Nominal PluralityWhich nouns take a plural, and is it always required?All nouns, always obligatoryAll nouns, always obligatoryAll nouns, always obligatoryAll nouns, always obligatoryAll nouns, always obligatory––All nouns, always obligatory
37A Definite ArticlesHow does the language say 'the'?Definite affixDefinite word distinct from demonstrativeDefinite word distinct from demonstrativeDefinite word distinct from demonstrativeDefinite word distinct from demonstrativeDefinite affixDefinite word distinct from demonstrativeDefinite word distinct from demonstrative
38A Indefinite ArticlesHow does the language say 'a'?Indefinite word same as 'one'Indefinite word same as 'one'Indefinite word same as 'one'Indefinite word same as 'one'Indefinite word same as 'one'Indefinite word distinct from 'one'–Indefinite word same as 'one'
45A Politeness Distinctions in PronounsDoes the word for 'you' change with politeness?Multiple politeness distinctionsBinary politeness distinctionBinary politeness distinctionBinary politeness distinctionBinary politeness distinctionNo politeness distinction–Binary politeness distinction
46A Indefinite PronounsHow are words like 'someone' and 'something' formed?Interrogative-basedGeneric-noun-basedGeneric-noun-basedSpecialMixed–Interrogative-basedInterrogative-based
47A Intensifiers and Reflexive PronounsIs 'himself' in 'he did it himself' the same word as in 'he hurt himself'?DifferentiatedDifferentiatedDifferentiatedDifferentiatedDifferentiated–DifferentiatedDifferentiated
49A Number of CasesHow many grammatical cases do nouns have?2 casesNo morphological case-markingNo morphological case-markingNo morphological case-marking–4 casesNo morphological case-marking3 cases
50A Asymmetrical Case-MarkingDo pronouns and nouns mark case in the same way?Additive-quantitatively asymmetricalAdditive-quantitatively asymmetricalNo case-markingAdditive-quantitatively asymmetrical–Syncretism in relevant NP-typesAdditive-quantitatively asymmetricalSyncretism in relevant NP-types
51A Position of Case AffixesWhere does the language put case markers on the word?Case suffixesNo case affixes or adpositional cliticsPrepositional cliticsNo case affixes or adpositional cliticsNo case affixes or adpositional cliticsCase suffixesNo case affixes or adpositional cliticsCase suffixes
52A Comitatives and InstrumentalsIs 'with a friend' said like 'with a knife'?IdentityIdentityIdentity–IdentityIdentity–Identity
53A Ordinal NumeralsHow are 'first', 'second', 'third' formed?First, two-th, three-thFirst, second, three-thFirst, second, three-thFirst, second, three-thFirst, second, three-thFirst, two-th, three-thFirst, second, three-thFirst, second, three-th
54A Distributive NumeralsHow does the language say 'two each' or 'two by two'?Marked by preceding word–No distributive numeralsNo distributive numerals–Marked by preceding wordMarked by preceding wordMarked by preceding word
57A Position of Pronominal Possessive AffixesIs 'my' or 'your' attached to the noun as an affix?No possessive affixesNo possessive affixesNo possessive affixesNo possessive affixes–No possessive affixesNo possessive affixesNo possessive affixes
63A Noun Phrase ConjunctionIs 'and' the same word as 'with'?'And' different from 'with'–'And' different from 'with''And' different from 'with'–'And' different from 'with''And' different from 'with''And' different from 'with'
65A Perfective/Imperfective AspectDoes grammar tell finished actions from ongoing ones?Grammatical marking–Grammatical markingGrammatical markingGrammatical marking–Grammatical markingGrammatical marking
66A The Past TenseDoes the past tense separate recent from distant past?Present, no remoteness distinctions–Present, no remoteness distinctionsPresent, no remoteness distinctionsPresent, no remoteness distinctions–Present, no remoteness distinctionsPresent, no remoteness distinctions
67A The Future TenseIs the future shown by a verb ending, not a helper word?No inflectional future–Inflectional future existsInflectional future existsNo inflectional future–No inflectional futureNo inflectional future
68A The PerfectHow does the language say 'has eaten', 'has done'?No perfect–From possessiveFrom possessiveNo perfect–Other perfectFrom possessive
69A Position of Tense-Aspect AffixesAre tense and aspect added to the start or end of the verb?Tense-aspect suffixesTense-aspect suffixesTense-aspect suffixesTense-aspect suffixesTense-aspect suffixesTense-aspect suffixesTense-aspect suffixesTense-aspect suffixes
70A The Morphological ImperativeWhich 'you' forms (one or many) have their own command form?Second singularSecond singularSecond singularSecond singular and second pluralSecond singularSecond singularSecond singular and second pluralSecond singular and second plural
71A The ProhibitiveHow is a ban like 'Don't go!' formed?Special imperative + normal negativeSpecial imperative + normal negativeNormal imperative + normal negativeSpecial imperative + normal negativeSpecial imperative + normal negativeNormal imperative + special negativeSpecial imperative + special negativeSpecial imperative + special negative
72A Imperative-Hortative SystemsAre commands built like 'let us go' and 'let him go'?Maximal systemNeither type of systemNeither type of systemNeither type of systemNeither type of systemMaximal systemMaximal systemMaximal system
74A Situational PossibilityHow does the language say 'can' (able to, allowed to)?Verbal constructionsVerbal constructionsVerbal constructionsVerbal constructionsVerbal constructionsVerbal constructions–Verbal constructions
75A Epistemic PossibilityHow does the language say 'it might be so'?Verbal constructionsVerbal constructionsVerbal constructionsVerbal constructionsVerbal constructionsVerbal constructions–Verbal constructions
76A Overlap between Situational and Epistemic Modal MarkingDoes one form mean both 'can' and 'might'?Overlap for both possibility and necessityOverlap for both possibility and necessityOverlap for both possibility and necessityOverlap for both possibility and necessityOverlap for both possibility and necessityOverlap for either possibility or necessity–Overlap for both possibility and necessity
77A Semantic Distinctions of EvidentialityDoes grammar show how the speaker knows something?No grammatical evidentials–Indirect onlyNo grammatical evidentials–Indirect onlyDirect and indirectNo grammatical evidentials
78A Coding of EvidentialityHow is the source of information marked?No grammatical evidentials–Modal morphemeNo grammatical evidentials–Verbal affix or cliticPart of the tense systemNo grammatical evidentials
81A Order of Subject, Object and VerbIn what order do subject, object and verb come?SVOSVOSVOSVOSVOSVOSVONo dominant order
82A Order of Subject and VerbDoes the subject come before or after the verb?SVNo dominant orderSVNo dominant orderSVSVNo dominant orderNo dominant order
83A Order of Object and VerbDoes the object come before or after the verb?VOVOVOVOVOVOVOVO
85A Order of Adposition and Noun PhraseDo words like 'in', 'on' come before or after the noun?PrepositionsPrepositionsPrepositionsPrepositionsPrepositionsPrepositionsPrepositionsPrepositions
86A Order of Genitive and NounDoes the owner come before or after the thing owned?Noun-GenitiveNoun-GenitiveNoun-GenitiveNoun-GenitiveNoun-GenitiveNoun-GenitiveNo dominant orderNoun-Genitive
87A Order of Adjective and NounDoes 'red' come before or after 'house'?Noun-AdjectiveNoun-AdjectiveNoun-AdjectiveNoun-AdjectiveNoun-AdjectiveNoun-AdjectiveAdjective-NounAdjective-Noun
88A Order of Demonstrative and NounDoes 'this' come before or after the noun?MixedDemonstrative-NounDemonstrative-NounDemonstrative-NounDemonstrative-NounDemonstrative-NounDemonstrative-NounDemonstrative-Noun
89A Order of Numeral and NounDoes 'three' come before or after the noun?Numeral-NounNumeral-NounNumeral-NounNumeral-Noun–Numeral-NounNumeral-NounNumeral-Noun
90A Order of Relative Clause and NounDoes the relative clause ('who left') come before or after the noun?Noun-Relative clauseNoun-Relative clauseNoun-Relative clauseNoun-Relative clauseNoun-Relative clauseNoun-Relative clauseNoun-Relative clauseNoun-Relative clause
90C Postnominal relative clausesWhich kinds of relative clause can follow the noun?Noun-Relative clause (NRel) dominantNoun-Relative clause (NRel) dominantNoun-Relative clause (NRel) dominantNoun-Relative clause (NRel) dominantNoun-Relative clause (NRel) dominantNoun-Relative clause (NRel) dominantNoun-Relative clause (NRel) dominantNoun-Relative clause (NRel) dominant
91A Order of Degree Word and AdjectiveDoes 'very' come before or after the adjective?Degree word-AdjectiveDegree word-AdjectiveDegree word-AdjectiveDegree word-AdjectiveDegree word-AdjectiveDegree word-AdjectiveDegree word-AdjectiveDegree word-Adjective
92A Position of Polar Question ParticlesIs there a yes/no question particle, and where does it go?No question particleNo question particleInitialNo question particleInitialInitialSecond positionInitial
93A Position of Interrogative Phrases in Content QuestionsIn 'What did you see?', does 'what' come first?Initial interrogative phrase–Initial interrogative phraseInitial interrogative phrase––Initial interrogative phraseInitial interrogative phrase
94A Order of Adverbial Subordinator and ClauseDoes 'because' come before or after its clause?Initial subordinator wordInitial subordinator wordInitial subordinator wordInitial subordinator word–Initial subordinator wordInitial subordinator wordInitial subordinator word
95A Relationship between the Order of Object and Verb and the Order of Adposition and Noun PhraseDoes verb–object order go together with preposition order?VO and PrepositionsVO and PrepositionsVO and PrepositionsVO and PrepositionsVO and PrepositionsVO and PrepositionsVO and PrepositionsVO and Prepositions
96A Relationship between the Order of Object and Verb and the Order of Relative Clause and NounDoes verb–object order go together with relative-clause order?VO and NRelVO and NRelVO and NRelVO and NRelVO and NRelVO and NRelVO and NRelVO and NRel
97A Relationship between the Order of Object and Verb and the Order of Adjective and NounDoes verb–object order go together with adjective order?VO and NAdjVO and NAdjVO and NAdjVO and NAdjVO and NAdjVO and NAdjVO and AdjNVO and AdjN
112A Negative MorphemesHow is 'not' expressed: affix, particle or verb?Negative particleNegative particleNegative particleNegative particleNegative particleNegative particleNegative particleNegative particle
115A Negative Indefinite Pronouns and Predicate NegationDoes 'nobody came' also need 'not' on the verb?Predicate negation also presentMixed behaviourMixed behaviourMixed behaviourMixed behaviour–Predicate negation also presentPredicate negation also present
116A Polar QuestionsHow are yes/no questions marked?Interrogative intonation onlyInterrogative intonation onlyQuestion particleInterrogative word orderQuestion particleQuestion particleQuestion particleQuestion particle
117A Predicative PossessionHow does the language say 'I have a book'?'Have'––'Have'–'Have'–'Have'
118A Predicative AdjectivesIn 'the house is big', does 'big' behave like a verb?Nonverbal encoding–Nonverbal encodingNonverbal encoding–Nonverbal encodingNonverbal encodingNonverbal encoding
119A Nominal and Locational PredicationAre 'he is a doctor' and 'he is in the house' built the same way?Identical–IdenticalDifferent–IdenticalIdenticalIdentical
120A Zero Copula for Predicate NominalsCan 'is' be left out in 'he is a doctor'?Impossible–ImpossibleImpossible–ImpossibleImpossibleImpossible
124A 'Want' Complement SubjectsIn 'I want to leave', is the subject of 'leave' said again?Subject is expressed overtly–Subject is left implicitSubject is left implicit–Subject is expressed overtlySubject is expressed overtlySubject is expressed overtly
143A Order of Negative Morpheme and VerbDoes 'not' come before or after the verb?NegVNegVOptDoubleNegNegVNegVNegVNegVNegV
143E Preverbal Negative MorphemesWhich negative words or prefixes come before the verb?NegVNegVNegVNegVNegVNegVNegVNegV
143F Postverbal Negative MorphemesWhich negative words or suffixes come after the verb?NoneNoneVNegNoneNoneNoneNoneNone
143G Minor morphological means of signaling negationIs negation also shown by tone or by changing the verb itself?NoneNoneNoneNoneNoneNoneNoneNone
144A Position of Negative Word With Respect to Subject, Object, and VerbWhere does the 'not' word go among subject, verb and object?SNegVOSNegVOOptDoubleNegSNegVOSNegVOSNegVOSNegVOMore than one position
144B Position of negative words relative to beginning and end of clause and with respect to adjacency to verbIs the 'not' word right before the verb, or elsewhere in the clause?Immed preverbalImmed preverbalImmed postverbalImmed preverbalImmed preverbalImmed preverbalImmed preverbalImmed preverbal
144D The Position of Negative Morphemes in SVO LanguagesIn SVO languages, where does 'not' go?SNegVOSNegVOOptNegSNegVOSNegVOSNegVOSNegVOMore than one construction
144H NegSVO OrderWhen 'not' opens the clause, is a second negative needed?No NegSVONo NegSVONo NegSVONo NegSVONo NegSVONo NegSVONo NegSVONo NegSVO
144I SNegVO OrderWhen 'not' follows the subject, is a second negative needed?Word&NoDoubleNegWord&NoDoubleNegWord&OnlyWithAnotherNegWord&NoDoubleNegWord&NoDoubleNegWord&NoDoubleNegWord&NoDoubleNegWord&NoDoubleNeg
144J SVNegO OrderWhen 'not' follows the verb, is a second negative needed?No SVNegONo SVNegOWord&OptDoubleNegNo SVNegONo SVNegONo SVNegONo SVNegONo SVNegO
144K SVONeg OrderWhen 'not' comes at the end, is a second negative needed?No SVONegNo SVONegNo SVONegNo SVONegNo SVONegNo SVONegNo SVONegNo SVONeg

Sounds: two descriptions of Romanian, two lists

PHOIBLE collects sound inventories. An inventory is the list of a language's sounds that can tell two words apart (its phonemes, sometimes with their variants), as a linguist has established it. There are 3,020 inventories for 2,186 languages. What is measured here is how many sounds two inventories have in common, out of all the sounds they have between them; one variant of the calculation measures instead how close the sounds of one inventory are to those of the other. It does not measure how often each sound occurs in speech, nor the melody of a sentence, nor how the language sounds to someone who does not know it.

For Romanian, PHOIBLE has three inventories, and they do not agree. Two follow the usual description: 7 vowels, 20 consonants and the semivowels of “iarnă” (winter), “ziuă” (day), “seară” (evening) and “soare” (sun). The third, compiled from the linguist Ioana Chitoran's 2002 book on the sound system of Romanian, counts separately the softened (palatalised) consonants at the end of words such as “lupi” (wolves), “pomi” (trees), “rupi” (you break): the p of “lupi” is entered as a different sound from the p of “lup” (wolf). The inventory thus reaches 51 segments, against 31 and 32.

Romanian's sounds in PHOIBLE's three descriptions

SPA (Agard 1958, Ruhlen 1973) · 31

consonants (20) bddʒgefɡghhkclmnprsʃșttsțtʃcevzʒj

vowels (7) aeəăiɨîou

semivowels and diphthongs (4) ji (iarnă)wu (ziuă)ɛ̯ɔ̯

UPSID (Agard 1958, Tătaru 1978, Ruhlen 1973) · 32

consonants (20) bddʒgefɡghhkclmnprsʃșttsțtʃcevzʒj

vowels (7) aeəăiɨîou

semivowels and diphthongs (5) ji (iarnă)wu (ziuă)e̞ae̞o̞o̞a

EA (Chitoran 2002) · 51

consonants (40) bbʲçddʲdʒgedʒʲffʲɡgɡʲhhkckʲllʲmmʲnnʲppʲrrʲssʲʃșʃʲttʲtsțtsʲtʃcetʃʲvvʲzzʲʒjʒʲ

vowels (7) aeəăiɨîou

semivowels and diphthongs (4) ji (iarnă)wu (ziuă)e̯äo̯ä

Filled: the sounds present in all three descriptions. Outlined: those only that description has. Diphthongs and the semivowels of “seară” and “soare” are written differently from one source to the next and are left out of the comparison.

With the usual description, the nearest inventories belong to distant languages: Damana in Colombia, Bangolan in Cameroon, Nupe in Nigeria, a French-based creole of Brazil (Karipúna) and Mazanderani in Iran. The first three have no connection to Romanian; Mazanderani is an Iranian language, so a very distant relative. Like Romanian, they have a medium-sized inventory made mostly of sounds found all over the world. Bangolan has all 29 sounds present in the three descriptions of Romanian: seven vowels, 20 consonants and the semivowels of “iarnă” and “ziuă”.

When the softened consonants are counted separately, the list moves to northern and eastern Europe: Russian, Sorbian, Veps, Skolt Saami, Lithuanian, Erzya, Bulgarian. All have whole series of softened consonants, as Russian does in “брать” (brat', to take) against “брат” (brat, brother).

With 7 vowels and 20 consonants

  1. Damana (Malayo) Colombia2 (1–11)
  2. Bangolan Cameroon2 (1–22)
  3. Nupe Nigeria6 (3–11)
  4. Karipúna French Creole Brazil9 (2–17)
  5. Mazanderani Iran10 (1–56)
  6. Sayula Popoluca Mexico16 (8–32)
  7. Ndamba Tanzania17 (4–37)
  8. Nigerian Pidgin Nigeria17 (3–54)
  9. Ishkashimi Afghanistan and Tajikistan20 (8–46)
  10. Kirundi Burundi21 (8–52)
  11. Antiguan and Barbudan Creole Caribbean23 (5–147)
  12. Dazaga Chad and Niger29 (10–52)
  13. Lunda Angola, Zambia, DR Congo29 (16–52)
  14. Mwani Mozambique30 (3–90)
  15. Yamba Cameroon38 (17–78)

With the softened consonants of Romanian “lupi” (wolves) and “pomi” (trees) counted separately

  1. Russian4 (1–5)
  2. Upper Sorbian Germany5 (3–7)
  3. Veps Russia5 (1–7)
  4. Skolt Saami Finland and Russia5 (1–10)
  5. Lithuanian5 (1–12)
  6. Votic Russia8 (6–10)
  7. Erzya Russia8 (4–11)
  8. Lower Sorbian Germany9 (6–9)
  9. Bulgarian10 (1–62)
  10. Ter Saami Russia12 (10–21)
  11. South Saami Norway and Sweden14 (1–17)
  12. Rusyn the Carpathians16 (9–27)
  13. Albanian17 (14–18)
  14. Komi-Zyrian Russia23 (15–64)
  15. Lealao Chinantec Mexico27 (10–46)

The median rank (the middle one when the variants of the calculation are put in order) and, in brackets, the best and the worst rank across the variants that use that description.

In none of the 72 variants of the calculation does Italian climb above place 19, and Mandarin Chinese never climbs above place 986, of 2,085. Albanian has 28 of the 29 sounds common to the descriptions of Romanian, Italian 25 (it lacks “ă”, “î”, “h” and the “j” of “joc”), Bulgarian 24, Mandarin 17.

Of those 29 sounds, the rarest in PHOIBLE's inventories are the “j” of “joc” (game), the sound of the s in “measure” (ʒ, present in 15.9% of inventories), “î” (ɨ, 16.5%) and “ă” (ə, 22.4%). Then come “ț” (ts, 26.8%) and “v” (27.0%). The most widespread sound missing from all three descriptions of Romanian is the “ng” of English “sing” (ŋ, in 62.8% of inventories). Among the next are the “ñ” of Spanish (ɲ, 42.7%) and the glottal stop, the catch in the middle of “uh-oh” (ʔ, 37.5%).

How widespread Romanian's sounds are

  1. ʒ j16%
  2. ɨ î17%
  3. ə ă22%
  4. ts ț27%
  5. v27%
  6. dʒ ge27%
  7. z34%
  8. ʃ ș37%
  9. tʃ ce41%
  10. f44%
  11. h h56%
  12. ɡ g57%
  13. d62%
  14. b63%
  15. e70%
  16. o70%
  17. r75%
  18. s77%
  19. l81%
  20. w u (ziuă)83%
  21. p86%
  22. u88%
  23. a90%
  24. j i (iarnă)90%
  25. k c91%
  26. t92%
  27. i92%
  28. m96%
  29. n97%
  1. ŋ63%
  2. ɲ43%
  3. ɛ38%
  4. ʔ37%
  5. ɔ36%
  6. iː32%
  7. aː31%
  8. uː30%
Each bar shows the share of PHOIBLE's 3,020 inventories that contain the sound.

How Romanian sounds to a stranger's ear is a different question. In an online game in which people guessed a language from a 20-second recording, with 15 million guesses analysed in 2017 by Hedvig Skirgård, Seán Roberts and Lars Yencken, on the map of the players' confusions Romanian sits next to Greek, Albanian and the Slavic languages.

Why the three answers do not agree

Each measure follows something else. The list of 40 words is made of the most stable meanings, and the first 17 places in the baseline calculation are all taken by Romance languages, Romanian's relatives.

Grammar is inherited, but it also moves towards the languages around it. Romanian is spoken far from the other Romance languages: Slavic languages and Hungarian lie between it and Italian, while Bulgarian, Albanian and Greek are nearer on the map. On the chart the Romance languages are near on both measures; Albanian, Greek and Bulgarian are near on grammar and far on words.

0.40.60.81.020%40%60%80%

← words: distance from Romanian↑ grammar: features with the same value

Each dot is a language with a word list and at least 40 grammar features described (243 languages). To the left: words nearer to Romanian's. Upwards: grammar nearer to Romanian's.

Sounds are few. A language has a few dozen phonemes, and the most widespread (p, t, k, m, n, a, i, u) turn up almost everywhere. Two medium-sized inventories can overlap by more than three quarters without the languages having anything to do with each other, as with Romanian and Damana.

That is why the question in the title has no single answer. “Closest” means something only together with the measure: on words, on grammar or on sounds.

How it was measured and which choices were made

Language identity

A language here is an entry at “language” level in the Glottolog 5.3 catalogue. Lists and inventories collected for dialects are assigned to their language. Standard Albanian rests on the Tosk dialect. ASJP's list for standard Albanian, attached to the group “Albanian”, and WALS's record “Albanian” (which has the ISO code sqi, no Glottolog code, and mentions both the Gheg and the Tosk dialect) were placed with Tosk Albanian, because PHOIBLE and Grambank keep their Albanian data under that code.

Words

ASJP version 21: 11,540 lists. The baseline calculation follows the project's rules: the 40-word list, at most two synonyms for a meaning, words marked as loans set aside, the Levenshtein distance adjusted for word length and divided by the average distance between words of different meanings (LDND). Romanian has three lists, which differ from one another by 0.12 to 0.22; the baseline is ROMANIAN_3. A language is ranked if it shares at least 28 words with the Romanian list: 5,457 languages. Left out: 405 lists with no Glottolog code, 22 lists of constructed languages, four lists attached to groups of languages and 620 languages with too few words.

The variants: the Romanian list (3), compound symbols taken as one unit or split (2), loans removed or kept (2), synonyms (the first two, all, or the closest pair: 3), the median list or the nearest when a language has several (2). 72 in all. The word table shows the median list; when that is the list of a local variety and ASJP also has a list named plainly after the language, the plain one is shown, with its own distance. For Mandarin Chinese, which has 184 lists, the table shows the Beijing list.

Grammar

WALS, edition 2020.4: 65 features of Romanian, after removing the 15 on phonology and the two on vocabulary; 28 of them describe word order. The baseline ranking requires at least 40 shared features: 243 languages, of the 1,900 languages that share at least one. The variants: all features, without the word-order ones (37), or in blocks, where features that follow from one another count once (51 blocks); with the threshold at 30, 40 or 50 shared features.

Grambank 1.0.3 was searched for Romanian's Glottolog code and for the whole Eastern Romance branch: the only record is Aromanian's. Aromanian was compared with the languages that share at least 120 features with it.

Sounds

PHOIBLE 2.0.1. Before comparison the notations were brought to one form: the marks for dental, apical, laminal, lowered, raised, centralised, advanced and retracted were removed; tapped r and trilled r were counted as one r; the semivowels written i̯ and u̯ were read as j and w. Tones, diphthongs and the semivowels of “seară” and “soare” are left out. Length, nasalisation, palatalisation, aspiration and voicing remain differences. This is a deliberate simplification: without it the same Romanian t would be three different sounds in the three descriptions.

The variants: the measure (overlap on the simplified notation, overlap on PHOIBLE's own notation, or a distance on PHOIBLE's 37 phonetic features: 3), diphthongs left out or kept (2), marginal phonemes kept or removed (2), the description of Romanian (3), the median inventory or the nearest when a language has several (2). 72 in all. In the sound-by-sound comparison a Romanian sound counts as shared if at least one inventory of the chosen language has it as an ordinary phoneme.

What is missing

No measure here says how easily two speakers understand each other, and none takes account of the whole vocabulary, of spelling or of actual pronunciation. The databases are uneven: WALS has few features for many languages, and PHOIBLE mixes descriptions made to different conventions. That is why every result is given together with the number of words or features it rests on.

The data, to download

The tables that contain PHOIBLE data (languages.csv and sounds-romanian.csv) are under CC BY-SA 3.0, like PHOIBLE; the others, derived from ASJP, WALS and Glottolog, are under CC BY 4.0.

Sources

The databases

Method and studies

On the same subject