Marius Comper

The science of learning

The narrow gate of memory

Someone reads you a phone number and you repeat it under your breath until you can dial it. If someone asks you the time in between, the number is gone. Almost everything you know, from the times tables to the way home, entered your mind through the same narrow gate: about four things at a time, held for roughly twenty seconds unless you keep repeating them. The best-tested rules of learning work with those two limits. Below you can measure your own gate and see what is lost from what has passed through it.

The test

Measure your gate

Digits appear one at a time, one second each. When they stop, you type them in order. The sequence grows by one digit after each correct answer, and the test ends on the second mistake. This is the classic digit-span task, used by psychologists for over a century; here it is a three-minute game rather than a clinical measure, and the result depends on tiredness, noise and how fast you rehearse.

How many digits you can hold at once

Press “Start” and watch the digits.

Miller’s famous figure, seven plus or minus two, comes from exactly this task. That does not mean the gate holds seven things: the digits of a phone number bind spontaneously into groups, and it is the groups that get counted. See below.

The same gate, with the door held open

You will see three letters, once. Then, for eighteen seconds, you count backwards in threes from a given number, typing each result. At the end you type the letters. This is the task with which Lloyd and Margaret Peterson measured in 1959 how long something survives in the mind when you are not allowed to rehearse it.

Press “Show the letters”.

100% 75% 50% 25% 0 0 s 3 s 6 s 9 s 12 s 15 s 18 s seconds of counting backwards before answering ≈ 80% ≈ 50% under 10%
Share of three-consonant groups recalled correctly, by how long the 24 students counted backwards before answering. The three values reported in reviews of the experiment are plotted: about 80% at 3 seconds, about 50% at 6 seconds, under 10% at 18 seconds. Source: Peterson and Peterson, Journal of Experimental Psychology, 1959.

The mechanism

Two rooms and a gate

The mind works with two very different spaces. The first is small and noisy: that is where you hold the phone number, the change you are owed, or the end of a sentence you are still reading. The second is vast and silent: that is where the language you speak, the faces you recognise and everything you ever learned are kept. Researchers call them working memory and long-term memory. Everything that reaches the second passed through the first, and the first is narrow.

How narrow? In 1956 the psychologist George Miller published a paper that became famous, “The magical number seven, plus or minus two”. The figure went into textbooks and office folklore. Miller himself later said that seven had been more of a rhetorical joke, a bridge between two unrelated lines of his research. When Nelson Cowan gathered, in 2001, dozens of experiments in which people could neither rehearse nor group what they saw, the real limit came out much lower: around four items in a typical adult. You get to seven with phone numbers because, without noticing, you bind the digits into groups.

How brief? That answer came from the experiment above. In 1959 the Petersons read students groups of three consonants, such as “CHJ”, and immediately had them count backwards in threes so they could not rehearse the letters. After three seconds, most remembered the group. After eighteen, fewer than one group in ten was still recalled. Without rehearsal, the gate empties in under half a minute.

In 1974 Alan Baddeley and Graham Hitch described this small room as a system with several parts: a loop that rehearses sounds, a sketchpad that holds images, and a conductor that divides attention between them. The model is still in use. What matters here is its consequence: working memory is less a box with four drawers than the ability to keep one pattern of neurons active while pushing away everything unrelated to it. When the pattern is reactivated often enough, the connections between its neurons strengthen and the pattern becomes easy to relight later. The process is called consolidation, and its result is long-term memory.

In short: a new thing has to be held active in the mind for several seconds running before it can start to consolidate. If it has more pieces than fit through the gate at once, it stays outside.

The trick

The secret is the size of the pieces

The gate counts chunks, meaning anything the mind already knows how to treat as a single thing, however many letters or digits a chunk contains. Here are thirteen letters.

CIAB BCFB INAS A

Thirteen letters, split at random: each group is a meaningless mark, far more than fits through the gate. Press the button to see what happens when the same letters are cut where you already know something.

The same letters, cut into CIA, BBC, FBI, NASA, become four chunks and pass through the gate in one go. The gate is unchanged; what changed is what you knew beforehand. Each chunk is a pattern already consolidated in long-term memory, and it lights up as a single item.

How far can this go? In 1980 Anders Ericsson, William Chase and a student, Steve Faloon, published the most spectacular answer. Faloon, a keen runner, started with an ordinary span of seven digits. After more than 230 hours of practice, he was repeating strings of 79 digits heard once. His working memory had stayed the same: he was turning digits into running times. “3492” became “3 minutes 49.2 seconds, close to the world record for the mile”. Given consonants instead of digits, he fell back to six or seven. The gate was as narrow as ever; his chunks were enormous, but only for digits.

This is where the first rule of effective learning comes from, the one every hastily written textbook trips over: a new thing must be taught after its parts have already become chunks. A child still counting on their fingers cannot hold a two-digit multiplication in mind, however well it is explained, because every addition in the middle of it takes up a place at the gate. The adult who knows the times tables by heart does it without any sense of effort, and the reason is that “7 times 8” is, for that adult, a single chunk.

Overload

When the gate jams

In 1988 John Sweller started from this limit and built an entire theory of teaching on it, cognitive load theory. The idea is simple. Any task requires holding a certain number of pieces in mind at once. If the number exceeds the gate, the task does not get solved and nothing is learned from it; if it merely approaches the limit, the task gets solved with difficulty and little is learned. A good explanation takes pieces out of the task rather than adding them.

Sweller and his colleagues collected dozens of measurable effects. One example: beginners learn more from an exercise already solved step by step than from the same exercise set as a problem, because free problem-solving fills the gate with random attempts. The effect has a reverse side, though, discovered by Slava Kalyuga in 2003. The same detailed explanations slow down an advanced student, who has to reconcile what they read with what they already know, and the reconciling also takes up places at the gate. The right help for someone who is just starting becomes noise for someone who has got it.

Children with a smaller gate

The differences between people are large and show early. Tracy and Ross Alloway followed 98 children in the United Kingdom from age five to eleven. Working-memory capacity measured at five predicted reading and mathematics scores at eleven better than IQ measured at the same time: working memory explained a part of the results that IQ did not. It is a single study with a small sample, but its conclusion fits what teachers see. Susan Gathercole, who spent years studying classrooms in Cambridge, describes children with poor working memory as “frequently inattentive”: they lose the thread of a three-step instruction, forget what they were supposed to write before the sentence is finished, seem not to be listening. Their gate is narrower, and they receive, like everyone else, instructions cut for a wider one.

The natural reaction would be to widen the gate. It has been tried on a large scale, with commercial “working-memory training” programmes. Two reviews led by Monica Melby-Lervåg and Charles Hulme, one in 2013 covering 23 studies and one in 2016 covering 145 experimental comparisons, reached the same result: training improves working-memory tasks similar to the ones practised, but the gain does not transfer to reading, arithmetic or reasoning. The gate is hard to widen. What can change is the size of the chunks and how they line up at the entrance.

The loss

What is lost after passing through

Passing through the gate is not the end of the story. A consolidated but unused pattern becomes harder and harder to relight. The first person to measure how fast that happens was a young German philosopher, Hermann Ebbinghaus, who in the 1880s memorised, alone, thousands of nonsense syllables and noted how long it took him to relearn them after ever longer intervals.

Photographic portrait of Hermann Ebbinghaus, a bearded man in a nineteenth-century suit
Hermann Ebbinghaus (1850–1909), the first to measure forgetting. Unknown photographer, before 1909. Source: Wikimedia Commons, Bettmann Archive. Public domain.

Ebbinghaus measured something finer than “how much you remember”: how much time you save when relearning. If a list took him ten minutes the first time and six minutes when he took it up again after a day, the saving is 40%. A saving of zero means the list had to be learned from scratch. The curve he obtained drops steeply in the first hour and then ever more gently. In 2015 Jaap Murre and Joeri Dros repeated the experiment to the letter, with a single volunteer who spent 70 hours learning lists by the original method. The two curves, 130 years apart, have the same shape: a steep drop in the first hour, then a slow one. The 2015 volunteer retained more at one and two days than Ebbinghaus did, and almost nothing at 31 days.

The shape of the curve says two things at once. The greatest forgetting happens in the first hours. And what survives the first day is then lost much more slowly.

60% 40% 20% 0 20 minutes 1 hour 9 hours 1 day 2 days 31 days time since learning (logarithmic scale) Ebbinghaus, 1880 2015 replication

Ebbinghaus, 1880 (published 1885)Murre and Dros, 2015

Time saved when relearning a list of syllables, as a percentage of the original learning time. Ebbinghaus: 58% after 20 minutes, 44% after one hour, 33% after 9 hours, 27% after one day, 23% after 2 days, 9% after 31 days. Replication (one volunteer): 56%, 47%, 34%, 46%, 43%, 4%. The replication’s authors themselves flag the 31-day value as unusually low compared with every other series. Source: Murre and Dros, PLOS One, 2015.

What works

Three habits that contradict intuition

If long-term memory strengthens only when the pattern is relit, then the best exercise is the relighting itself: pulling the information out of your mind, unaided, instead of pushing it in once more through your eyes. The experiments that tested this produce some of the largest differences in the whole psychology of learning.

Recall it instead of rereading it

In 2006 Henry Roediger and Jeffrey Karpicke gave students short scientific texts. One group had four five-minute reading sessions, in which it went through the text about 14 times on average. The other had a single reading session, about three passes, and spent the other three sessions writing down from memory everything they could recall, without looking at the text. Five minutes after the last session, the readers were doing better and were convinced they would remember more. A week later the order had reversed: the readers reproduced 40% of the text, those who had tested themselves 61%. The group that had spent the most time with the text in front of them had forgotten it fastest.

Two years later the same authors asked what happens after you have already learned something. Students memorised 40 Swahili–English word pairs, testing themselves until they knew them all. Then half kept testing themselves on the words they already knew, while the other half dropped them from testing, as anyone does with “learned” flashcards. After a week, the first group remembered about 80% of the pairs, the second 36% or 33%, depending on how much they had reread. Rereading words already known left the result unchanged; testing them more than doubled it.

Scientific texts, reproduced after a week

one reading session, three test sessions61%
four reading sessions40%

Roediger and Karpicke, Psychological Science, 2006, experiment 2.

Swahili word pairs, after a week

still tested after being known≈ 80%
dropped from testing, reread36%
dropped altogether33%

Karpicke and Roediger, Science, 2008.

Biology text, questions after a week

reproduced from memory81%
concept map with the text open58%
reread in sessions57%
read once50%

Karpicke and Blunt, Science, 2011, experiment 1. The students had predicted that rereading would help them most.

Leave time between repetitions

The second rule concerns the calendar. The largest experiment over long intervals was run by Nicholas Cepeda, Harold Pashler and colleagues in 2008, with 1,354 volunteers recruited online. Each learned 32 small facts until they knew them perfectly, reviewed them once after a pause set by the researchers, and was tested at a deadline of up to a year after the review. Study time was equated across groups; for the same deadline, only the pause differed. The result was always hump-shaped: too short a pause does not help, the right pause helps enormously, and too long a pause helps somewhat less. At the right pause, those tested remembered on average 64% more than those who had reviewed immediately.

When to review, according to this experiment

How far the test is from the review

If the test comes 7 days after the review, the best pause between learning and review in the experiment was one day.

These are the four deadlines actually measured, counted from the review to the test, with the pauses that gave the best results among those tried: 1, 11, 21 and 21 days. For other deadlines the experiment measured nothing, so this page invents nothing. The authors’ rule of thumb: the further away the deadline, the longer the best pause in days but the shorter as a share of the deadline, from nearly a fifth for a few weeks to under a tenth for a year.

Mix the exercises

The third rule is the hardest to believe. Doug Rohrer and Kelli Taylor taught students to calculate the volumes of four little-known solids. Everyone solved the same problems; only the order differed. One group got them in blocks, all of one kind, then all of another, as in an ordinary textbook. The other got them mixed. During practice, the blocked group did better, because it knew in advance which formula was coming. On the test a week later, the mixed group scored 63% and the blocked group 20%. Mixing forced each student to choose the method every time, and choosing is exactly what a test demands.

The effect appears even where there is no formula. Nate Kornell and Robert Bjork showed students paintings by twelve artists, either grouped by artist or mixed, then asked them to attribute new paintings. Mixing won, but most participants declared with confidence that they had learned better from the grouping. A 2026 rerun of this experiment found the same direction of effect again, more modest in size: 38% correct after mixing against 29% after grouping.

Geometry problems, test after a week

mixed practice63%
blocked practice20%

Rohrer and Taylor, Instructional Science, 2007.

Painters’ styles, new paintings to attribute

paintings mixed38%
paintings grouped by artist29%

Direct replication of Kornell and Bjork (2008), published in 2026.

The illusion

Why it feels exactly backwards

In several of the experiments above, the participants predicted wrongly what would help them. Those who reread felt better prepared. Those who solved in blocks solved faster during practice. Those who saw the paintings grouped were convinced they had learned better. The feeling of ease is real; it just measures something else.

Ease tells you how open the gate is right now. Rereading, grouping and immediate repetition keep the pattern lit in working memory, so everything feels familiar and clear. But that very familiarity is the sign that long-term memory has not been put to work. Robert Bjork called the habits that help “desirable difficulties”: the strain of pulling something out of your mind, of choosing the method, of returning after a pause in which you have had time to forget a little. In the experiments above, the version that felt harder at the time was, every time, the one that was still there a week later.

What a landmark review says

In 2013 John Dunlosky and four colleagues went through all the available literature on ten learning techniques commonly used by pupils and students and graded them by how broad and how solid their effect is. The result contradicts the most widespread study habits.

Ten learning techniques and their utility according to the 2013 Dunlosky review
TechniqueWhat it meansUtility
Practice testingPulling the information out of your mind, with flashcards, questions, free recall.high
Distributed practiceSpreading the same study hours over days and weeks.high
Interleaved practiceAlternating problem types instead of doing them in blocks.moderate
Elaborative interrogationExplaining to yourself why a fact is true.moderate
Self-explanationSaying out loud what you are doing and why, as you solve.moderate
RereadingGoing through the text again.low
HighlightingMarking the important passages.low
SummarisationWriting a summary of the text.low
Keyword mnemonicLinking a term to a similar-sounding word and an image.low
Imagery for textImagining scenes for each paragraph.low

Source: Dunlosky, Rawson, Marsh, Nathan and Willingham, Psychological Science in the Public Interest, 2013. “Low” means the effect is weak, hard to reproduce or limited to certain subjects; it does not mean the technique does harm.

The whole mechanism, in brief

Four rules derived from two limits

  1. Teach after the parts have become chunks

    A new thing passes through the gate only if its parts are already consolidated and light up as a single item. The order of the material is part of teaching.

  2. Cut everything into pieces that fit

    If a step requires more than a few things held in mind at once, it is not learned. The worked example helps the beginner and can hinder the advanced student.

  3. Pull out rather than push in

    Long-term memory strengthens when you relight the pattern from inside. Self-testing beats rereading by margins of tens of points, including for things you already know.

  4. Let yourself forget a little, then return

    The same study hours, spread over weeks and mixed together, yield more than in blocks. It usually feels harder precisely when it is working.

What remains uncertain

The edges of these figures

The figure of four items is an average of adults under laboratory conditions, and the individual value ranges from about three to about five; the exact number depends on what is counted and how grouping is prevented. The eighteen seconds come from one experiment with 24 students and nonsense syllables; with meaningful material the duration differs, and part of the rapid loss is explained by interference between successive trials rather than by the passage of time alone. The study linking working memory at five to results at eleven has 98 children and, as far as I could find, has not been repeated at the same scale; the wording “better than IQ” is the authors’ and holds for that sample. The 2015 forgetting curve comes from a single volunteer, and the authors themselves flag his 31-day value as unusually low. The mixing effect with painters came out more modest in recent reruns than in the original experiment. The “best” pauses in the planner above are the best among those tried in that experiment, not a mathematical optimum. Nothing on this page is medical or educational advice for a particular case.

Method and sources

Where the figures come from

This page brings together results published in peer-reviewed journals, read in their original form wherever possible. The two tests at the top reproduce the format of the classic experiments, with randomly generated digits shown for one second each; their result is a game, not a standardised test. The charts reproduce the values from the cited papers’ tables, without any processing; the review planner shows only the four deadlines measured in the 2008 experiment. The 31-day point of the 2015 replication is kept as published, together with the authors’ warning.

  • George A. Miller, “The magical number seven, plus or minus two”, Psychological Review, 1956. psycnet.apa.org
  • Nelson Cowan, “The magical number 4 in short-term memory: A reconsideration of mental storage capacity”, Behavioral and Brain Sciences, 2001. pubmed.ncbi.nlm.nih.gov
  • Lloyd R. Peterson and Margaret J. Peterson, “Short-term retention of individual verbal items”, Journal of Experimental Psychology, 1959. pubmed.ncbi.nlm.nih.gov
  • Alan Baddeley and Graham Hitch, “Working memory”, in The Psychology of Learning and Motivation, vol. 8, 1974. doi.org
  • K. Anders Ericsson, William G. Chase and Steve Faloon, “Acquisition of a memory skill”, Science, 1980. doi.org
  • John Sweller, “Cognitive load during problem solving: Effects on learning”, Cognitive Science, 1988. doi.org
  • Slava Kalyuga, Paul Ayres, Paul Chandler and John Sweller, “The expertise reversal effect”, Educational Psychologist, 2003. doi.org
  • Tracy Packiam Alloway and Ross G. Alloway, “Investigating the predictive roles of working memory and IQ in academic attainment”, Journal of Experimental Child Psychology, 2010. pubmed.ncbi.nlm.nih.gov
  • Susan Gathercole, “Understanding how working memory problems impair classroom learning”, MRC Cognition and Brain Sciences Unit, Cambridge. mrc-cbu.cam.ac.uk
  • Monica Melby-Lervåg and Charles Hulme, “Is working memory training effective? A meta-analytic review”, Developmental Psychology, 2013; Melby-Lervåg, Redick and Hulme, “Working memory training does not improve performance on measures of intelligence or other measures of far transfer”, Perspectives on Psychological Science, 2016. pubmed.ncbi.nlm.nih.gov
  • Jaap M. J. Murre and Joeri Dros, “Replication and analysis of Ebbinghaus’ forgetting curve”, PLOS One, 2015, table 3. journals.plos.org
  • Henry L. Roediger III and Jeffrey D. Karpicke, “Test-enhanced learning: Taking memory tests improves long-term retention”, Psychological Science, 2006. pubmed.ncbi.nlm.nih.gov
  • Jeffrey D. Karpicke and Henry L. Roediger III, “The critical importance of retrieval for learning”, Science, 2008. doi.org
  • Jeffrey D. Karpicke and Janell R. Blunt, “Retrieval practice produces more learning than elaborative studying with concept mapping”, Science, 2011. pubmed.ncbi.nlm.nih.gov
  • Nicholas J. Cepeda, Edward Vul, Doug Rohrer, John T. Wixted and Harold Pashler, “Spacing effects in learning: A temporal ridgeline of optimal retention”, Psychological Science, 2008. doi.org
  • Doug Rohrer and Kelli Taylor, “The shuffling of mathematics problems improves learning”, Instructional Science, 2007; the figures are restated in Rohrer, “Interleaving helps students distinguish among similar concepts”, Educational Psychology Review, 2012. doi.org
  • Nate Kornell and Robert A. Bjork, “Learning concepts and categories: Is spacing the ‘enemy of induction’?”, Psychological Science, 2008; direct forced-choice replication, 2026. pubmed.ncbi.nlm.nih.gov
  • John Dunlosky, Katherine A. Rawson, Elizabeth J. Marsh, Mitchell J. Nathan and Daniel T. Willingham, “Improving students’ learning with effective learning techniques”, Psychological Science in the Public Interest, 2013. doi.org