mariuscomper.ukRomână

Dan Ungureanu and the origin of Romanian

A linguist says Romanian is the Latin of northern Italy, spoken after AD 500. Is he right?

The Romanian linguist Dan Ungureanu reads the history of his language differently from the textbooks. His idea fits in one sentence with three parts. Each part can be checked on words Romanians say every day.

  • holds up on the data
  • promising, not yet proven
  • open or contradicted
1

When

When did Romanian split from the rest of Latin?

What Dan Ungureanu saysRomanian went through the same vocabulary changes as the spoken Latin of the West until late, around 540–600. It did not break away early, before 250, as an isolated "Danubian Latin" would require.

His evidence is a medieval study aid. The Reichenau Glossary, written in northern France, explains old Latin words that readers no longer understood with new words from spoken Latin. For example femur, thigh, is glossed with coxa. And Romanian says coapsă: it kept the new word.

Ungureanu took 66 such pairs and counted which side Romanian is on. All of them are below. Blue: Romanian kept the new word. White: it kept the old one.

See for yourself

49
took the new word, like seară from sera
17
kept the old word, like caș, cheese, from caseum, where French says fromage

All 66 pairs Ungureanu counted.

  • coapsăfemur → coxa
  • potârnichecoturnices → quacoles
  • aramăaes → aeramen
  • amenințacogor → anetsor
  • bășicăpapula → vesica
  • sarcinăsarcina → bisatia
  • turtăcolliridam → turtam
  • împrumutamutuare → impruntare
  • tăiaconcidit → taliavit
  • brâutorax → brunia
  • ficatjecore → ficatum
  • îngânadeludere → degannare
  • plângefletur → planctur
  • dezlegasolveris → disligaveris
  • arinăarena → sabulum
  • furafurent → involent
  • foaleutres → folli
  • măsuraremetietur → remensurabit
  • purtasublatum → subportatum
  • nastureinstitis → nasculis
  • fagurefavum → frata mellis
  • treieratereo → tribulo
  • lutluto → fecis
  • schimbacommutatione → concambiis
  • ariearea → danea
  • cașcaseum → formaticum
  • călcâicalx → calcaneum
  • descoperidenudare → discooperire
  • teacăfaretra → teca
  • gliegleba → blicta
  • oaieoves → berbices
  • pușchepustula → malis clavis
  • îngreunaoneratus → graviatus
  • aprindesuccendunt → sprendunt
  • umblapergite → ambulate
  • legejus → legem
  • înecasubmersi → necata
  • înjunghiajugulate → occidite
  • știnosse → scire
  • ascundeabdito → absconso
  • searăvesper → sera
  • bunătatebenignitas → bonitas
  • dada → dona
  • juneadolescentia → juventus
  • puteanequeo → non possum
  • cetateoppidum → civitas
  • stânglevam → sinistramnot counted
  • cumpăraemo → comparo
  • mascurmas → masculus
  • câmpager → campus
  • totcuncti / universi / omnes → toti
  • tememetuo → timeo
  • suflaflare → sufflare
  • desdensus → spissus
  • încărcaoneratus → carcatus
  • cuvenioportet → convenit
  • intraingredi → intrare
  • ugerubera → mamilla
  • tăceasileo → taceo
  • negruater → niger
  • fierbedifferveo → exbullio
  • crăpăturăfissura → crepatura
  • ușăfores → ostia
  • iarnăhiems → hibernus
  • jucaludo → ioco
  • șoarecemus → sorex
  • auăuva → racemus

On each card: old word → new word from the glossary, with the Romanian word above. In bold: the Latin word Romanian kept. The dotted card, stâng, is shown by Ungureanu but left out of the count.

How sure is it?

The count holds. The date needs more evidence.

  • The count reproduces. Re-read row by row, his coding still gives 49 out of 66, though some links between the Latin and the Romanian word remain debatable (the "Only the secure links" filter).
  • But it does not give a year by itself. To get from "74% new words" to "540–600" you need to know when each new word spread. Some were already old: niger and campus appear in Cicero and Virgil.
  • The other Romance languages also mostly took the new words. On 46 of the pairs, coded provisionally from reference works, French has the most, as expected of a glossary written in France, and Sardinian, which split early, is above Romanian. This does not refute Ungureanu, but it shows what is needed: a date for each word.
French93%
Catalan91%
Occitan89%
Romansh89%
Sardinian77%
Romanian73%
Italian72%
Spanish63%
Portuguese59%

The share counts only pairs where the language continues one of the two words. Provisional coding from reference works, on 46 pairs.

Deeper: the numbers behind the date

Ungureanu proposes four calculation scenarios, giving 530, 540, 540 and 600. His formula cannot be rebuilt from a single rule, so 540–600 is the output of his model, not a statistical confidence interval.

If it were pure luck, with independent pairs and equal odds for the two words, 49 of 66 would be almost impossible (p ≈ 0.00005). The result stays significant until you expect, from the start, 64% of pairs to come out "new". So the open question is: what should you expect?

The full study, card by card, with the dictionaries consulted (in Romanian): Reichenau, la firul cuvântului. Ungureanu's paper: "Glosarul din Reichenau și limba română", Journal of Humanistic and Social Studies 15/2 (2024), pp. 105–117.

2

Where from

Which Latin does Romanian come from?

What Dan Ungureanu saysThe Latin Romanian grew out of was not just any Latin. It was most like the dialects of Liguria, western Lombardy and the Alps, from Dolomite Ladin to Alpine Occitan. Since 2024 he says the language then took shape south of the Danube, along the Roman road called the Via Egnatia.

His finest piece of evidence is a sound. Romanian turned a Latin l between two vowels into r: mola became moară, mill; gula became gură, mouth; solem became soare, sun. Standard Italian does not do this. But poems written in Genoa around 1300 do exactly the same.

Genoa around 1300, and Romanian today

LatinItalianGenoese, c. 1300Romanian
gula throatgolagoragură mouth
molamolamoremoară mill
angelusangeloangeroînger angel
solemsolesorsoare sun
gelugelozerger frost
qualisqualequarcare which
caelumcielocercer sky
NicolausNicolosoNicherosoNicoară

Latin gula meant "throat": Genoese kept the meaning, Romanian moved it to the mouth. The Genoese forms come from the Rime genovesi of the late 13th century, published in Archivio Glottologico Italiano II (1876), as Ungureanu gathers them in his chapter on rhotacism in Istoria limbii române.

"ul vin e bun, ul vin e tulbure, ul vin e reo; le osa me doer; suna le campane"A sentence in Milanese that Ungureanu gives as an example: "the wine is good, the wine is cloudy, the wine is bad; my bones ache; the bells ring". In Romanian: vinul e bun, vinul e tulbure, vinul e rău; oasele mă dor; sună clopotele. Read aloud, a Romanian can follow it almost word for word.
Dan Ungureanu's map: the areas of France and Italy where l between vowels became r. Red: Liguria, Occitan Piedmont, Lombardy, Ticino, Ladinia, Provence; the dotted outline reaches Lyon and Burgundy.
Dan Ungureanu's map from Istoria limbii române (Cartier, 2024), labels in Romanian: the areas where l between vowels became r. Solid red: attested areas. Dotted outline: the greatest historical extent.

The route he proposes

Ungureanu does not say Romanians came straight from Lombardy. In his 2024 book, a Latin of the northern-Italian type reaches the Balkans, and Romanian takes shape south of the Danube, along the Via Egnatia, the Roman road that joined the Adriatic to Constantinople. The book names the area this Latin came from and the area where Romanian took shape. It does not describe the road between them, and no document records such a migration.

Sketch of Ungureanu's hypothesis: from northern Italy to the Balkans, then the Via EgnatiaMilanoGenovaDyrrachiumThessalonicaConstantinopleRomedialects with l → rVia EgnatiaDanube
The two areas in Ungureanu's hypothesis: red, the dialects Romanian most resembles; gold, the Via Egnatia, where he places the formation of Romanian.

How sure is it? Eight questions, in turn

Real likenesses. The exact place is not yet proven.

To say "Romanian comes from here", a likeness has to pass eight questions. Here is how far the strongest piece of evidence, l turning into r, has got:

  1. Does Romanian really have the feature?Yes: moară, gură, soare, cer, înger.
    yes
  2. Is it a shared change, not an old trait everyone kept?Yes: Latin had l; Romanian and the northern dialects both changed it.
    yes
  3. Is it found only in the proposed area?Mostly: Liguria, Lombardy, the Alps, south-eastern France. At Bronte, in Sicily, Ungureanu ties the r to northern settlers at nearby Maniace; other southern traces are uncertain.
    mostly
  4. How many independent clues are there?One. Eight words with r are one sound change, not nine clues.
    one
  5. Was it already there when this Latin left for the east?Unknown. The Genoese forms date from 1300, some seven centuries after the split he proposes.
    undated
  6. Common descent or later contact?Unresolved.
    open
  7. Who actually moved?No inscription or document yet shows a large departure from north-western Italy.
    open
  8. Where did Romanian take shape?The Via Egnatia is his 2024 proposal; it needs its own test.
    open

Step 4 is the lesson most easily forgotten: however long the list, one sound change counts once. That is why Ungureanu gathers other features too: doi/două, two, like Gallo-Italian dui/due; eu sunt, I am, beside Po-valley io sonto; zănatic, crazy, from dianaticus, with a parallel in Turin.

Deeper: the critics, and village by village

Ioana Vintilă-Rădulescu (Studii și cercetări lingvistice 69/1, 2018) objected that he does not show the features to be exclusive to northern Italy. Rudolf Windisch (Zeitschrift für Balkanologie 60/2, 2024) finds the method acceptable. Dan Alexe objects that it ignores the Balkan language area, and Alexandru Nicolae (HotNews interview, 2024) doubts how far reconstructions of unattested periods can be verified.

The village-by-village test, on the linguistic atlas of Italy and 374 maps annotated by Ungureanu, is in the villages of Italy and Switzerland that say "masă". There, on his advice, only innovations count, not archaisms. It also shows that the exact form tulbur appears in only 11 Lombard villages, although the turbulus type is spread from Graubünden to Naples.

3

Before Latin

How many Dacian words does Romanian have?

What Dan Ungureanu says"Substrate words? Yes. Dacian words? No." A word Romanian shares with Albanian is not automatically Dacian. When you do not know which language a word comes from, the honest label is "unknown".

In 2019 he went through 189 old Romanian words with no secure Latin etymology, the ones usually called "substrate" words. For each he said where he thinks it comes from. The result:

1of 189 stays possibly Dacian in his list: brusture, burdock. And he calls even that one contested.

All 189. Tap a word

Tap a word to see the origin Ungureanu proposes.

The colour shows the origin Ungureanu proposes; when he gives two, the first counts. Framed words were checked in the dictionaries (DEX, DER, MDA). Words are spelled as in Ungureanu's catalogue, with î for â.

How sure is it?

The critique holds. His alternative origins, case by case.

  • The negative part is solid. The dictionaries do not prove any word Dacian either. Even brusture has a Slavic etymology in DER.
  • Some of his alternatives are confirmed by the dictionaries: traistă, bag, from Byzantine Greek tagistron; a îngâna, to mumble, from Latin ingannare; stâng, left, from *stancus; sterp, barren; gușă, goitre.
  • Others clash with them. DEX labels țap, billy goat, Thraco-Dacian; Ungureanu calls it Italic. The Celtic label for balegă, dung, is in none of the dictionaries checked.
  • The same bar for everyone. If "looks like Albanian" is not enough to make a word Dacian, "looks like Welsh" is not enough to make it Celtic.
Deeper: where the numbers come from

The catalogue: Dan Ungureanu, "Cît din substratul lexical al limbii române e dacic?", Analele Banatului (2019). Of 189 words he rejects 119 as Dacian, says a Dacian origin is not established for 69, and leaves one as possible. The 22 framed words were compared with DEX, DER and MDA for this page.

4

His method

How do you tell kinship from coincidence?

All three ideas above rest on one rule Ungureanu has repeated since 2010: before you say two words look alike, measure how often they would look alike by chance. The more "related" meanings you allow and the fewer sounds you require to match, the more false matches you find. He writes that only 2–3% of reconstructed roots are secure enough for very deep comparisons.

The coincidence machine

Two languages made up at random, 100 words each. They are not related at all. Choose how strictly you compare.

–

The same test, on the Altaic languages

Ungureanu compared the basic vocabulary of the Turkic, Mongolic and Manchu languages and found 22 matches in 600 comparisons. His conclusion: there is no proven Altaic family. But everything depends on how many matches you would expect by chance.

–

Ungureanu worked with 1%. At 1%, 22 matches are too many for chance, so his own arithmetic shows a signal. The result becomes compatible with chance only from about 2.5%. Specialists still disagree about Altaic.

He applies the same rule to reconstructions. Two dictionaries of the same African language family, Nilo-Saharan, reconstruct about 170 Proto-Nilo-Saharan forms (Bender) and more than 1,600 (Ehret). For Ungureanu, a gap that large calls for a statistical test. On its own, though, it does not prove the family does not exist.

Something rare in a researcher: in 2010–2011 Ungureanu argued for very old links between Indo-European, Uralic and Altaic. Applying the same rule about coincidences, he changed his conclusion. His method overturned his own belief.

5

Beyond language

Can a country have a personality?

In Zidul de aer (The Wall of Air, 2008), a treatise on mentalities, Ungureanu describes Romania as particularist, hierarchical and low in trust, and argues that a stable mentality undermines institutions. In his essays since 2017 he explains more and more through mechanisms you can see: roads, migration, incentives. Emigration, he wrote in 2023, replaced moving to the city and broke school as a social lift.

How individualist is Romania? It depends on the question. The same "individualism", which Ungureanu also quotes, measured on Romanians with different questionnaires:

0 = collectivist, 100 = individualist.

Scores for Romania: Hofstede, the 2005 Interact survey and three different dilemmas in the Trompenaars database. The 81 is the score Ungureanu quotes.

This spread does not refute his point about distrust, which recent OECD data also bear out. It shows how much a "national personality" depends on the instrument that measures it.

And a historian's rule: check the document first

The first mention of "Romanians"?

A medieval text cited as the first mention of the name rumeni reads, in Ungureanu's reading, cumeni, Cumans. Ungureanu showed the transcription error before discussing what it would mean.

The Rohonczi Codex

The enigmatic manuscript, read by some as a Dacian text, is written on Italian paper with a watermark from 1529–1540. Ungureanu started from the paper (Observator Cultural, 2003).

7

Books

What Dan Ungureanu has written

  • 2024Istoria limbii române (The History of the Romanian Language), Cartier, Chișinău. The synthesis of the ideas above.
  • 2016Româna și dialectele italiene (Romanian and the Italian Dialects), Romanian Academy Press. The first form of the northern-Italian hypothesis.
  • 2014Reconstructing Languages, with Alexandru Șișu, Romanian Academy Press. On coincidences and reconstructions.
  • 2011Relațiile lexicale între indo-europeană și familiile uralică și altaică: ipoteza nostratică (the Nostratic hypothesis), Romanian Academy Press.
  • 2008Zidul de aer. Tratat despre mentalități (The Wall of Air), Bastion.
  • 1999Originile grecești ale culturii europene (The Greek Origins of European Culture), Amarcord.
  • 2002Originea limbajului și primul om (The Origin of Language and the First Human), West University Press.

He is an associate professor at Aurel Vlaicu University in Arad. He has also written widely as an essayist, in Observator Cultural, CriticAtac and Vatra.

Sources

  1. Dan Ungureanu, "Glosarul din Reichenau și limba română", Journal of Humanistic and Social Studies 15/2 (2024), pp. 105–117.
  2. Dan Ungureanu, Istoria limbii române, Cartier, 2024, chapter on the rhotacism of intervocalic -L-.
  3. Dan Ungureanu, "The Geographic Origin of Romanian", Journal of Humanistic and Social Studies 7/1 (2016).
  4. Dan Ungureanu, "Cuvinte de substrat? Da. Cuvinte dacice? Nu.", Journal of Humanistic and Social Studies 8/1 (2017).
  5. Dan Ungureanu, "Cît din substratul lexical al limbii române e dacic?", Analele Banatului (2019).
  6. Dan Ungureanu, "Long-range comparisons and word roots decay", Journal of Humanistic and Social Studies 1/1 (2010).
  7. Dan Ungureanu, "There is no Altaic linguistic family", published on Academia.edu.
  8. Dan Ungureanu, Zidul de aer. Tratat despre mentalități, Bastion, 2008.
  9. Rime genovesi della fine del secolo XIII e del principio del XIV, Archivio Glottologico Italiano II (1876).
  10. Ioana Vintilă-Rădulescu, review of Româna și dialectele italiene, Studii și cercetări lingvistice 69/1 (2018), pp. 105–158.
  11. Rudolf Windisch, review of Istoria limbii române, Zeitschrift für Balkanologie 60/2 (2024), pp. 271–286.
  12. Cezar Bălășoiu, review of Istoria limbii române, Studii și cercetări lingvistice 75/1 (2024), pp. 127–133.
  13. Dan Alexe, "Câteva considerațiuni despre «Istoria limbii române» a lui Dan Ungureanu", Cabal in Kabul, 13 April 2024.
  14. Interview with Alexandru Nicolae, HotNews, 3 November 2024.
  15. Culture scores: F. Trompenaars, via the ASE Management Conference (2013); G. Hofstede; Interact Romania (2005).
  16. Dictionaries: DEX, DER (Cioranescu), MDA.

The page's data: the 66 pairs, the catalogue of 189 words, 165 claims by Ungureanu, each with its source.