mariuscomper.ukRomână

Language and mind, by the data

Bilingual children: are they smarter? What 149 studies show about attention and memory

From 1987 to 2020, 149 studies compared children growing up with two languages to children growing up with one, on six kinds of test, from ignoring distractors to working memory. All 1,194 results are here, one dot each. The average difference comes out small, and after correcting for selective publication no clear evidence of an advantage remains. Pick the children's age and the kind of test and see where the average falls.

Data: the public table of Lowe, Cho, Goldsmith and Morton (Psychological Science, 2021).

The short answer

The average advantage is small, and it is not clear that bilingualism causes it.

Across the 1,194 results pooled by Lowe and colleagues (2021), bilingual children come out slightly ahead of monolingual children: g = 0.08. The g value says by how many standard deviations two averages differ. Zero means no difference, and 0.2 is the threshold psychologists call "small"; 0.08 is below it. For illustration, assuming normal distributions with equal variance, a g of 0.08 puts the bilingual group's mean at the 53rd percentile of monolinguals. It is not a percentile observed in any child.

Positive results get published more easily than null ones. After the authors adjust for that, the estimate becomes −0.04: no clear evidence of an average advantage remains. For adults, a meta-analysis (Lehtonen and colleagues, 2018) and a study of 11,041 people (Nichols and colleagues, 2020) reached the same conclusion, and in 4,524 children aged 9 to 10, Dick and colleagues (2019) found little evidence of an advantage.

  • 0.08average difference in standard deviations, 95% confidence interval 0.01 to 0.14
  • −0.04after correcting for selective publication
  • 53percentile of monolinguals at which the bilingual mean would sit, for illustration

The result concerns laboratory tests in children aged 3 to 17. It says nothing about school grades, about what a child gains by being able to talk to grandparents in their language, or about any particular child.

Each dot is a result: try every filter

A dot is one comparison of bilingual and monolingual children on one test. To the right bilingual children scored higher, to the left monolingual children. The higher the dot, the more children in the comparison. The vertical line at zero means no difference between bilingual and monolingual children.

Children's age
What is measured

All six kinds of test together.

Where it appeared

Not published in a journal: doctoral theses, manuscripts and unpublished data.

Studies newer than the chosen year appear faded.

An illustrative scenario: this is how the literature would look if only what came out "significant" had reached print.

  • significant in favour of bilinguals
  • significant in favour of monolinguals
  • not distinguishable from zero
  • model average, with confidence interval
  • beyond the line, a result comes out significant (approximately)

What to notice: the low dots, with few children, scatter widely, because with 20 to 50 children a large difference also arises by chance. The high dots, with hundreds or thousands of children, cluster closer to zero, with exceptions.

Loading the results…

Averages by category, on all results

The same averages as a table, for anyone who cannot use the chart. The model is refitted from the authors' code, on their public table. The categories are grouped from the table's tests and do not coincide with the domains in the authors' paper: their "executive attention" domain has g = 0.06 (interval −0.02 to 0.14). The age rows do not add up to the total: 44 results have no age recorded.

Filterg95% intervalResultsGroups of children
All results0.0750.013 … 0.1381,188189
Children under 50.130−0.025 … 0.28619438
Children aged 5–70.055−0.053 … 0.16342668
Children aged 8 and over0.0770.002 … 0.15252474
Ignoring distractors0.029−0.071 … 0.12937193
Switching rules0.049−0.065 … 0.16227078
Working memory0.052−0.047 … 0.151331101
Attention0.152−0.049 … 0.35310519
Stopping a response0.1740.047 … 0.3025730
Other0.060−0.037 … 0.1585417
Only in scientific journals0.0770.011 … 0.1441,101174
Not published in a journal0.036−0.118 … 0.1908715

The ten largest studies

Ordered by the largest comparison in each. Authors and years are as in the public table. "Average of results" is the simple average of the study's results.

StudyMost childrenAverage of resultsResults
Dick 20194,5240.003
Santillan 2018949−0.071
Lesaux 2003938−0.313
Choi 2018685−0.101
Soliman 20146120.6911
Brito 2018534−0.043
Dunabeitia 2014504−0.036
Anton 2014360−0.0914
Jaekel 20193370.124
Arizmendi 2018247−0.214

How the pooled estimate changed as studies accumulated

The cumulative average over studies published up to a given year: for 2005 (55 results from 20 groups of children) it was 0.145, and with all studies it reaches 0.075, with fluctuations in between. The interval narrowed at the same time: from 2016 it no longer includes zero, but the average stayed small. The values describe the evidence base as it grew, not a steady decline, and mix tests, research groups and ages.

−0.10.00.10.20.30.4g, cumulative averagegroups200520200830201045201260201491201612220181702020189−0.04 after correction
The hollow diamond shows the value the authors published after correcting for selective publication: −0.04.
The values in the chart
Up tog95% intervalResultsGroups of children
20050.145−0.037 … 0.3265520
20080.143−0.064 … 0.3509730
20100.126−0.033 … 0.28525245
20120.118−0.010 … 0.24537660
20140.084−0.008 … 0.17661291
20160.1140.030 … 0.198817122
20180.0740.008 … 0.1391,117170
20200.0750.013 … 0.1381,188189

Why a small effect can look big

In a study of 50 children (two equal groups), a difference has to exceed about 0.55 standard deviations to come out "significant". Small studies have little power to detect a difference of 0.08: their estimates jump around from one study to the next, and a large difference can also arise by chance. A large result gets published more easily; a null one reaches readers with more difficulty.

The 219 results significant in favour of bilinguals (by the test described under "How the numbers are calculated"; 18% of all) have a simple average of 0.98. The model average over all results is 0.08. The other tail has 71 results (6%) significant in favour of monolinguals. If every underlying effect were zero, the tests well calibrated and nothing selected, about 2.5% would be expected in each tail. Heterogeneity and selection can both raise these percentages, so this cannot show whether the imbalance comes from a small real effect or from selection.

The unweighted average is 0.28 in the third of comparisons with the fewest children (17–50 children), 0.18 in the middle third and 0.02 in the third with the most (66–4,524 children). The comparison does not adjust for differences between tests or for several results from the same children.

The share of results significant in favour of bilinguals does not fall clearly with study size: 20% in the third with the fewest children, 16% in the third with the most (the thirds hold 460, 343 and 391 results). Among the largest studies there are exceptions too, as in the table above.

Researchers have tracked selection directly. De Bruin and colleagues (2015) analysed conference abstracts from 1999 to 2012 and checked which were later published. Using the classification reproduced in Leivada's table (2023), 27 of 40 abstracts with results favouring bilinguals were published (68%), against 5 of 17 with null or negative results (29%). Leivada disputes part of the classification. The comparison suggests selective publication under that coding; it does not say how much of the pooled average selection explains.

Adults, old age and the language you learn in

For adults, Lehtonen and colleagues (2018) pooled 891 results from 152 studies. Before correction they found a very small advantage in three of six domains. After correcting for selective publication no advantage remained. For verbal fluency (how many words you produce in a minute) a small disadvantage appeared. Nichols and colleagues (2020), with 11,041 people and twelve tests, found a bilingual advantage on one test, while monolinguals did better on four. All the differences vanished once the groups were matched to remove potentially confounding factors.

On dementia, studies disagree because they measure different things. Craik and colleagues (2010) analysed 211 consecutive patients diagnosed with Alzheimer's disease: the bilingual patients had been diagnosed 4.3 years later and had reported first symptoms 5.1 years later, and the groups were equivalent on cognitive and occupational level (the monolinguals had more schooling). But these are patients who reached a clinic, and Mukadam and colleagues (2017) note that retrospective studies are more prone to confounding by education or by cultural differences in how people reach dementia services. In prospective studies they found an odds ratio of 0.96 (interval 0.74–1.23, 5,527 bilingual participants) and found no protection against cognitive decline or dementia.

Vocabulary, assessing young children, and school

In children, Bialystok and colleagues (2010) analysed 1,738 children aged 3–10 and found a consistently smaller receptive vocabulary in bilinguals, in the language they were tested in, mostly for words from home life. Dick and colleagues (2019) found the same for English, and the gap shrank substantially once socioeconomic status or intelligence was taken into account.

In young children, Core and colleagues (2013) compared 47 Spanish–English bilingual children aged 22 to 30 months with 56 monolingual children. Total vocabulary, meaning the words from both languages together, had means and growth rates similar to the monolinguals' and identified the same proportion of children below the 25th percentile. The authors recommend total vocabulary for assessing language in young bilingual children. The study concerns assessment, in one sample; it does not say what to do if a child speaks very little. Genesee and colleagues (1995) followed five bilingual children aged 1 year 10 months to 2 years 2 months: they mixed words from their two languages, but clearly told them apart.

Mother tongue clearly matters elsewhere: at school. According to the World Bank (2021), about 37% of students in low- and middle-income countries are required to learn in a language different from the one they speak. The World Bank writes that children learn more and are more likely to stay in school if they are first taught in a language they speak and understand. It is a problem of how schools are organised, not a comparison between languages.

Other claims about mother tongue

The same question, asked of other claims in circulation. Each row links to the study itself.

ClaimWhat was foundVerdict
"Russians tell shades of blue apart faster, because their language has two words for it."Winawer and colleagues (2007): 26 Russian and 24 English speakers, 21 of each left in the analysis; with close shades, Russians were faster when the shades fell on either side of the boundary between "goluboy" and "siniy", and the advantage vanished when speakers silently rehearsed a long number. Martinovic and colleagues (2020), in two experiments, did not find the advantage at the same boundary. Winawer et al. 2007 Martinovic, Paramei & MacInnes 2020Mixed evidence for a language-linked difference in speed. These tasks alone do not establish a change in how you see.
"Grammatical gender changes how we see objects."Samuel and colleagues (2019) reviewed 43 papers with 5,895 participants: the effect shows in some tasks and not in others. A preregistered replication with 375 participants (Elpers and colleagues, 2022) did not find it for natural languages. Samuel, Cole & Eacott 2019 Elpers et al. 2022Weak and task-dependent.
"Languages without an obligatory future make people save more."Chen (2013) reported an association across 76 countries: speakers of languages where the future is optional were 31% more likely to have saved in a year. Roberts, Winters and Chen (2015) took language relatedness into account: the association is weaker and, in a model with all survey waves (logit coefficient 0.26, interval −0.06 to 0.57), no longer significant; the authors note, though, that it stayed reasonably robust under other tests. Chen 2013 Roberts, Winters & Chen 2015A weak association, with no evidence that grammar causes it.
"Transparent number words make children better at maths."Miller and colleagues (1995): 99 Chinese and 98 American children aged 3 to 5. Significant differences in counting appeared at ages 4 and 5, not at 3, and only in the teens decade: up to 20, 74% of the Chinese children counted and 48% of the American; up to 10, 92% and 94%. Between 20 and 99, as on simple problems, there were no differences. Lê and Noël (2020), with 104 Vietnamese and 104 French-speaking Belgian children aged 3½ to 5½, found the gap again only in counting. Miller, Smith, Zhu & Zhang 1995 Lê & Noël 2020Real and narrow: a head start in counting at 4 to 5 and not at 3, in the names from 11 to 19, which cannot separate language from family and early education.
"Your mother tongue decides how intelligent you are."Intelligence test scores rise by 1 to 5 points for each year of schooling (Ritchie and Tucker-Drob, 2018), so a comparison between languages cannot separate language from schooling, income and health. No validated instrument for wisdom by language is known. Ritchie & Tucker-Drob 2018It cannot be established from existing data.
"Some languages are more efficient: they say more in the same time."Coupé and colleagues (2019), 17 languages, computed from their data: information per second ranges from 33.8 to 45.9 bits, although fast speakers use syllables that carry less. Romanian was not measured. Coupé, Oh, Dediu & Pellegrino 2019Among the 17 languages measured, the highest rate is at most 1.36 times the lowest; fast speakers say less per syllable. No ranking of "efficiency" follows.

On this site

How the numbers are calculated

  • The data are the public table of Lowe and colleagues, from OSF: 1,194 results from 149 sources (136 journal articles, 11 doctoral theses and two unpublished data sets), 189 independent groups of children and 98 research groups, with their coding and direction of calculation (a positive g means an advantage for bilinguals) kept unchanged.
  • The specification in the authors' code was refitted in Python: the average of the results, with nested random effects (research group, study, group of children), known sampling variances for each result, REML estimation and a normal-approximation interval. On all results it reaches g = 0.075 (interval 0.013 to 0.138), against the published 0.08 (0.01 to 0.14) on the same 1,188 results. Every filter on the page refits the model; with fewer than eight results or four groups of children no average is shown.
  • As in the authors' code, the 6 results with |g| above 3 are left out of the averages. The chart's axis is narrower than the extreme values: dots beyond the edge are counted under the chart.
  • The colours show whether each result's interval, computed with a normal approximation and unadjusted, excludes zero: z above 1.96 favours bilinguals, below −1.96 favours monolinguals. They may differ from the tests in the original papers. The dotted curves are approximations for equal groups. Simple averages in the text are labelled as such.
  • The correction for selective publication was not recomputed: −0.04 is the authors' value. The averages on the chart are not corrected.

What the data do not say

  • The children are aged 3 to 17, and about 51% of the results come from the USA and Canada.
  • These are laboratory tests. There are no school grades, jobs or happiness here.
  • A group average says nothing about any particular child.
  • Mother tongue and bilingualism are different things: the chart compares children with and without a second language, not languages with one another. The table of other claims reports studies that compare groups, and none supports a ranking of languages.
  • These are the studies the authors gathered. There is no guarantee they are all that exist.

Data and code, so you can check

Everything the page computed can be rebuilt. The pack below holds the 322 filter averages, the figures on the page with their sources, the code and a script that downloads the authors' table, checks its fingerprint and refits the model.

From the pack folder:

pip install -r requirements.txt
cd code
python3 reproduce.py

It should print g = 0.075 (interval 0.013 to 0.138), against the published 0.08 (0.01 to 0.14) on the same 1,188 results.

The authors' table is not in the pack: the OSF project that holds it declares no licence (their preprint is CC0), so the pack points to the source. The page draws its dots from their numbers, as published, with credit.

Licences: the averages and numbers computed here, CC BY 4.0; the code, MIT. The cited studies and their abstracts remain their authors'.

Found a mistake or a number that does not reproduce? Write to bilingual@mariuscomper.uk.