mariuscomper.uk Română

After Dan Ungureanu, “Istoria limbii române” (A History of the Romanian Language), Cartier, 2024

Romanian words in the dialects of Italy: a map of Dan Ungureanu’s “Istoria limbii române”

Romanian is the Romance language of Romania and Moldova. The usual account places the Latin it descends from in or near the Roman province of Dacia, in today’s Romania, or south of the Danube; the Romanian linguist Dan Ungureanu argues that it was spoken in north-western Italy. For 525 Romanian words, his book sets forms from the dialects of Italy, France and Switzerland beside them, with the region and sometimes the village. This map puts each form where the book puts it. Among the regions of Italy, Lombardy has the most words, then Veneto and Sardinia. The north of Italy has a form for 319 words; the centre, the south and the islands for 284.

The book’s whole vocabulary at a glance

The number in each circle shows how many words have, in the book, at least one form from that region. North of the La Spezia–Rimini line, which separates the dialects of northern Italy from the rest, 319 words have a form; south of it or on the islands, 284. Of these, 130 have forms on both sides.

The map shows only where the forms cited in the book are found. That Romanian comes from north-western Italy is the author’s thesis. Of two reviews, one calls his method acceptable and the other calls it deeply mistaken.

A check: on a list of 110 basic words, Romanian shares a root with each Romance language on the list outside the Balkans in between 69 and 80 cases. No word is shared with Piedmont alone or with Liguria alone. The list has no French and no Lombard dialect. See the check

  • Or pick one:
  • (frost)
  • (to tickle)
  • (mouth)
  • (cloudy)
  • (only)
  • (mill)

Aosta ValleyFrench-speaking SwitzerlandPiedmontTicinoLombardyGraubündenTrentino-Alto AdigeLiguriaSouthern FranceFranco-Provençal areaIstriaEmilia-RomagnaMarcheVenetoFriuliNorthern FranceAbruzzoMoliseApuliaSalentoBasilicataCalabriaCampaniaLazioTuscanySicilySardiniaCorsicaUmbriaLadin areaLa Spezia–Rimini line155Lombardy123Veneto87Northern France85Sardinia83Ladin area80Liguria68Tuscany61Southern France58Sicily58Emilia-Romagna56Salento53Calabria52Lazio4640373028272524242219191511411
numbered circle: how many words have a form in the region
region named in the book
form given for two regions together
town or area named in the book
form from an old text, placed at its author’s town
town the book calls a Ligurian or Gallo-Italic colony

Blue: what the book says. Red: what the linguistic atlas of Italy and southern Switzerland (AIS) recorded in the 1920s. The book’s data and the atlas’s are shown separately. Base map: Natural Earth.

Among the regions of Italy, Lombardy, Veneto and then Sardinia have the most words

The chapter “Vocabular selectiv” (selective vocabulary) has 679 entries for 649 distinct headwords. For 581 of them the book gives at least one form from another language or dialect; 525 have a form within the area the map covers. The chapter’s stated aim is to show that Romanian comes from a late spoken Latin, not to collect words found only in the north; in some entries the book itself says the form is found throughout Italy or in the south.

Among the regions of Italy, Lombardy has the most words (155), followed by Veneto (123) and Sardinia (85). The Ladin area has 83, Liguria 80. Piedmont, which the thesis includes in the area of origin, has 19 in this chapter. Taken together, forms from France (Old and modern French, Occitan, Franco-Provençal) appear for 152 words.

The La Spezia–Rimini line, drawn between the two towns from the Ligurian Sea to the Adriatic, is the conventional boundary between the dialects of northern Italy and the rest. North of it 319 words have a form; south of it or on the islands, 284. Of these, 130 are on both sides, 189 only in the north and 154 only in the south or on the islands. Of the 189 found only in the north, 45 also have, in the book, a form from France or the Iberian Peninsula. The north-west, that is Piedmont, the Aosta Valley, Liguria, Lombardy and Ticino, has 212 words; 76 have forms in Italy only there.

north only: 189 · both sides: 130 · south or islands only: 154

The book calls a few towns in the south and on the islands Ligurian or Gallo-Italic colonies: Sassari and Sant’Antioco in Sardinia, Bonifacio in Corsica, Bronte and San Fratello in Sicily. Forms cited from them are shown on the map as points and are not counted for the south: 33 words have such forms. Tuscan words the book gives only as a contrast (“vs. tosc.”) are not counted either, nor are the seven matches it calls separate loans or coincidence.

Two things limit the count. The vocabulary is the author’s choice: it contains, as he writes, only words that are “interesante fonetic, morfologic sau semantic” (interesting in sound, form or meaning), so the numbers describe the book, not the Romanian language. And a region with many dialect dictionaries and many old authors is more likely to be cited.

How it was counted: a word counts for a region if the book gives at least one form attributed to that region. A form that comes with the name of a place identified on the map counts for that place’s region. A form that comes only with the label of a dialect counts for the region or regions of that label (“calabr.-sicil.” means Calabria and Sicily). The label “ital.” counts for no region; “Italia de nord” (northern Italy), with no region, counts for the north. The north: Piedmont, Aosta Valley, Liguria, Lombardy, Ticino (the Italian-speaking canton of Switzerland), Emilia-Romagna, Veneto, Trentino-Alto Adige, the Ladin area, Friuli and Istria. Regions are counted whole: the north of Marche, which lies beyond the line, goes with the centre. The labels are the book’s and do not follow regional borders: “lomb.” takes in villages in Ticino and Piedmont, “ligur” takes in Monaco, “romand” means Franco-Provençal in a wide sense, as far as Grenoble and Aosta, and “ladin” (Ladin, a Romance language of the Dolomite valleys) takes in Cadore and Val di Sole.

  • 155
  • 123
  • 87
  • 85
  • 83
  • 80
  • 68
  • 61
  • 58
  • 58
  • 56
  • 53
  • 52
  • 46
  • 40
  • 37
  • 30
  • 28
  • 27
  • 25
  • 24
  • 24
  • 22
  • 19
  • 19
  • 15
  • 11
  • 4
  • 3
  • 1
  • 1
The number of words in the book with at least one form in the region. Solid: the north. Grey: the centre, the south and the islands. Outline only: Graubünden, French-speaking Switzerland and France. Select a name to see the words.

Can a list of words show where a language comes from?

Romanian and the dialects of Italy descend from the same language, Latin, so many shared words are to be expected. A more precise place of origin can be argued only from shared, exclusive innovations: changes Romanian shares with one area and with no other. The German linguist Karl Brugmann stated the rule in 1884. A Latin word kept in two places and lost in a third shows nothing, however many such words are collected.

The author himself writes that he counts only innovations, and Rudolf Windisch’s review notes the same. The checks below, made on the linguistic atlas of Italy and southern Switzerland (AIS, compiled by Karl Jaberg and Jakob Jud from fieldwork in the 1920s) and on northern Italian texts of the 12th and 13th centuries, show how hard it is for any one word to meet that condition.

A word list not chosen for its resemblances

On Saenko’s list of 110 basic meanings (2015), Romanian shares a root with:

Aromanian94
Lower Engadine Romansh80
standard Italian, Neapolitan79
Friulian, Fassa Ladin, Piedmontese (Barbania)78
Venetian76–77
Piedmontese (three other points)75–77
Ligurian71–75
Spanish75
Sardinian73–75
Catalan69–70

Aromanian, spoken in the Balkans, is Romanian’s closest relative. Piedmontese and Neapolitan score almost the same: 75–78 and 79. Matches Romanian has with one group only: two with the Iberian Peninsula, one with the Rhaeto-Romance area of the Alps, one with Sardinia, none with Piedmont or Liguria. The list includes no Lombard dialect and no French, so it cannot test the very region with the most words in the book.

Three examples the author put forward in conversation, checked village by village

Not all three words are from the vocabulary: “după” has no entry in it, and for “numai” and “vreun” the book itself gives, in other chapters, the forms outside northern Italy listed below. Villages are classified by rules on the root of the word, not by a dialectologist.

WordWhere it is south of the Alps (AIS atlas)ElsewhereWhat it shows
numai (only)
non magis
55 of 378 villages, all north of the lineCatalan només, Asturian namái, Spanish nomás, Old French nemésa late Latin expression that survived; Tuscany and the south lost it
vreun (any, some)
vere unus
32 villages: 23 in eastern and Alpine Lombardy, 4 in Graubünden, 4 in Trentino, 1 in UmbriaAromanian vãrnu; Old Italian has it in Tuscany tooan old Italian word that fell out of use almost everywhere
după (after)
de post
the dop-, dup- type in 305 of 379 villages, in every regionFrench depuis, Spanish después, Portuguese depoisa word found across the Romance languages

The one cluster that was found

Forms of the “numai” type (Romanian for “only”, from Latin “non magis”) and the change of l to r occur together in 12 villages, nine of them in Ticino. Neighbouring villages are not independent evidence, so the 12 count as one piece of evidence, for one area. Ten of them have the form with d- (“doma”), whose link to “non magis” is assumed, not demonstrated. Chance alone would give about five shared villages, not 12. The same check keeps, as innovations worth following up, the change of l to r, the metathesis “tulbur” (11 villages), the feminine gender of “miere” (honey) and “fiere” (bile), and “gula” with the meaning “mouth”.

The atlas has a limit of its own: it shows what was left after more than six centuries in which Tuscan replaced local forms. The poet Bonvesin da la Riva, writing in Milan around 1280, has both “verun” (any) and “dra” (Tuscan “della”), with r for l; in the atlas of the 1920s the two features no longer occur together in any village.

What the book argues and how it was received

In “Istoria limbii române” (Cartier, Chișinău, 2024), Dan Ungureanu argues that Romanian descends from a late Latin spoken after the year 350 in north-western Italy, in an area he bounds by Milan, Lugano, Grenoble and Genoa, from which its speakers would later have left for the Balkans. The usual account places that Latin in or near the Roman province of Dacia, in today’s Romania, or south of the Danube. As evidence the book offers the changes that, according to the author, Romanian shares with the dialects of that part of Italy: l turned into r (“soare”, sun; “moară”, mill); metatheses, sounds that have swapped places, as in “tulbure” (cloudy), from Latin “turbulus”, and “întreg” (whole), from “integru”; and new meanings such as “gură” (mouth), from Latin “gula”, throat.

In the preface the author reconstructs a sentence as it might have sounded in northern Italy between 600 and 700. It begins “Eo ao beut vin tulbure” (I have drunk cloudy wine); in Tuscan the same words would have been “Eo ho bevuto vino torbido”. It is the author’s reconstruction, not a surviving text.

Two reviews appeared in 2024. The Romance scholar Rudolf Windisch, in “Zeitschrift für Balkanologie” (60/2), summarises the book chapter by chapter and calls the comparison of Romanian forms with those of northern Italy and beyond the Alps “eine akzeptable Methode”, an acceptable method; he also notes that the book’s three overlaid maps are “kaum erkennbar”, barely discernible. Cezar Bălășoiu, of the University of Bucharest, in “Studii și cercetări lingvistice” (LXXV/1), concludes that the book’s method is “profund eronată”, deeply mistaken, because it starts from the premise that Dacia was completely abandoned at the end of the third century, and objects, among other things, that it never says which linguistic features the La Spezia–Rimini line rests on.

The objection goes beyond the geography of words. In the usual classification, which Bălășoiu recalls, the line separates Western Romance (Portuguese, Spanish, Catalan, French, Occitan, Rhaeto-Romance and the dialects of northern Italy) from Eastern Romance, which takes in the dialects of central and southern Italy, Dalmatian and Romanian. One of the features that define it is what happened to voiceless consonants between vowels: Latin “rota” (wheel) gave “roată” in Romanian and “ruota” in Italian, but “rueda” in Spanish and “roue” in French.

The thesis in full, with the author’s replies to his critics, is on the page about Dan Ungureanu and in the dossier of 165 claims. The atlas village by village, for 19 words: The dialects of Italy.

The preface table: the forms the author sets side by side

The book’s preface sets the Latin word, the Romanian word and a dialect form side by side. Below are 12 rows of that table, with the place the author gives. The dotted line on the map is the La Spezia–Rimini line, drawn here after the map in the book.

LatinRomanianDialect formWhere, according to the book
turbulustulbure (cloudy)tulbureLombard
integruîntreg (whole)intreghLombard
monte, bonu, sonaremunte, bun, a suna (mountain, good, to ring)munt, bun, sunaLombard
exvolarea zbura (to fly)sbioraBonifacio (Corsica), a Ligurian colony, according to the preface
ranabroască (frog)broscoa Lombard word
essea fi (to be)firall of northern Italy
malurău (bad)reonorthern Italy
alvearealbină (bee)alvinaFriuli
gula “throat”gură (mouth)guraonly in the Alps
labiumbuză (lip)bozaPragelato
—a dîrdîi (to shiver)derdelarLombardy and France
—a gîdila (to tickle)gatiłarnorthern Italy and France

Two of these rows were checked in old texts. “Reo” is also literary Tuscan, so it is not confined to the north; and in Bonvesin “brosco” means a toad, where Romanian “broască” is a frog.

All 649 words of the vocabulary

The words are spelled as in the book, with î inside the word (“cîmp”, “a gîdila”). In grey, those for which the book gives no form from a Romance dialect.

How the ledger was made and what is not shown here

The ledger, that is the list of forms taken from the book, was extracted automatically from the manuscript in two independent passes: one by rules (the region label, then the form printed in italics), the other by a language model reading each paragraph. The first produced 2,165 rows, each with one region and one form, the second 2,249; they agreed on 1,808, including 55 rows where one pass also put the form under an extra label. The 820 cases on which the passes disagreed were checked against the text in a third pass: 485 accepted, 335 rejected. Every form was found letter for letter in its paragraph. The ledger has 2,269 rows; one accepted case can give several rows, when the same form stands under two labels. Apart from eight rows (four keyed by hand from the preface table and four added in the third pass), no row entered the ledger on the strength of a single pass. None of the passes was done by a dialectologist.

In an independent check of 40 words, 374 rows were compared with the book’s text: 15 were wrong, most of them with the place attached to the neighbouring form, and no form the book clearly gives was missing. Since that check, a place receives a point only if it is printed right beside the form. Errors certainly remain. No paragraph of the chapter was left unattached to an entry.

On the map, 185 places have a point. A town or village received a point only if the GeoNames list or the list of AIS atlas villages has exactly one place of that name in the region the book indicates; the other names stay in the ledger without a point. Valleys, old provinces and departments have their point at their centre, and old texts in their author’s town, from tables drawn up by hand. The Ladin area on the map is the outline drawn around the places the book itself calls Ladin.

The chapter refers 64 times to numbered maps of the AIS atlas. Eight references were set aside because the title given in the book is missing or does not match the title of the map with that number. For the rest, villages and forms come from the atlas’s digital edition.

The author’s text, his examples from old texts and the book’s images are not reproduced here. The page gives only the data of each entry: the word, the Latin word cited, the region, the place and the form. The book’s synthesis maps are by Paul Șteț and Ligia Pop; the one that overlays phonetics, morphology and vocabulary was drawn by Ligia Cornelia Pop under the supervision of Professor Alina Satmari.

Sources

  • Dan Ungureanu, Istoria limbii române, Chișinău, Cartier, 2024, ISBN 978-9975-86-757-3. The chapter “Vocabular selectiv” and the preface table.
  • Rudolf Windisch, review (in German), Zeitschrift für Balkanologie 60 (2024), no. 2, pp. 271–286.
  • Cezar Bălășoiu, review (in Romanian), Studii și cercetări lingvistice LXXV (2024), no. 1, pp. 127–133.
  • Michele Loporcaro et al., AIS, reloaded, v2.0.0, doi:10.5281/zenodo.19832160, licence CC BY-NC-SA 4.0: the villages, coordinates and forms of Jaberg and Jud’s atlas.
  • Saenko, “Annotated Swadesh wordlists for the Romance group”, 2015, in the Global Lexicostatistical Database (CC BY 4.0).
  • Alexandre François and Siva Kalyan, “Subgrouping: Trees vs. waves”, 2022; Daniel Kaufman, “Subgrouping”, 2026: the rule of shared innovations.
  • Natural Earth, scale 1:10,000,000 (public domain): coasts, regions, provinces. GeoNames (CC BY 4.0): town coordinates.

Read also: The Latin Island: the Romanian language · Milk from lactem