# What remains of psychology

The marshmallow test, Stanford prison, power poses: what can we still trust in stories about the human mind? Follow 48 ideas from their original promise to the evidence that came later.

Some findings held up. Others became narrower or were contradicted. The difference lies in the question tested and the evidence for each case.

Revised 5 September 2026

## A marshmallow cannot decide a child’s future

Waiting for a reward was associated with later school achievement. In a more diverse sample, that relationship shrank substantially after accounting for family background and early abilities. The developmental question remains interesting; a test lasting a few minutes calls for care when used to judge a child.

[What changed in the marshmallow test](https://mariuscomper.uk/ce-a-ramas-din-psihologie/en/#marshmallow-test)

## What were the “guards” told to do?

The Stanford prison study became a story about the power of roles. Later analysis of documents and recordings showed how the researchers shaped the situation. Understanding the behavior also requires reading the instructions participants received.

[Read the Stanford case](https://mariuscomper.uk/ce-a-ramas-din-psihologie/en/#stanford-prison)

## Something you can use in your next lesson

Spread learning across sessions and try answering from memory before rereading. Research on spacing and retrieval practice supports these approaches. The useful interval depends on the material and how long you want to remember it.

[The evidence on learning](https://mariuscomper.uk/ce-a-ramas-din-psihologie/en/#spaced-repetition)

## When other researchers try again

In 2015, the Open Science Collaboration published replications of 100 studies from three psychology journals. 97 original results had been reported as significant. Of these, 35 crossed the statistical threshold again in the same direction: about 36%.

These 100 studies came from articles published in three journals in 2008. That selection and criterion cannot give a percentage of “true psychology.” The authors also compared effect sizes and the compatibility of results.

### What is a replication?

A new study tests the same question with different participants. Similarities in procedure, population, and measurement matter when interpreting the result.

### What does a p-value tell us?

How unusual a result at least this extreme would be if the null hypothesis and the test’s assumptions held. Effect size and precision answer different questions.

### What does a failed replication mean?

It weakens confidence in the claim tested. Statistical power, the procedure, and the wider evidence help distinguish a small effect from a context-dependent or absent one.

## How easily can you find a “significant” result?

You have an imaginary study with 30 participants. Exposure X and six outcomes Y were generated independently. Try other analyses on the same observations and watch the p-value. It can fall below 0.05 even though the simulated population contains no relationship.

Exploration can uncover worthwhile questions. The problem arises when the attempts disappear from the report and an analysis chosen after seeing the data is presented as a planned test. Reporting the full path and checking on new data make the difference.

The initial test uses Pearson correlation between X and the first outcome Y for the first 30 participants. Six outcomes and three control variables are generated before the first analysis. Adjustment computes a partial correlation, removing the linear relationship with one control variable. The two-sided p-value uses Student’s t distribution with n − 2 degrees of freedom, or n − 3 after adjustment. The plot shows residuals of both variables when an adjustment is used.

Exclusion tests every available observation, keeps the removal with the lowest p, and repeats once. The counter includes those searches and all three candidate adjustments. Displayed p-values are uncorrected for selection and repeated testing. The final check uses independent observations, the chosen outcome and adjustment, with no exclusions. After that check, the analysis ends; you can restart or generate another study. Restarting reuses all observations, including the checking sample; generate another study for new data.

[Try the lab](https://mariuscomper.uk/ce-a-ramas-din-psihologie/en/#phacker)

## 48 ideas, with the evidence in view

Find a story you know or choose a subject. Each case keeps the claim, subsequent research, and sources together. Its label applies to the specific claim described.

Supported (10): The cited research supports the phenomenon under the conditions described.

Supported with limits (22): Part of the claim holds; its size or generality needs qualification.

Unsupported as stated (12): The cited evidence does not support the specific promise described.

Contested (4): The cited findings or interpretations remain in disagreement.

### Ego Depletion

Roy Baumeister, Ellen Bratslavsky, Mark Muraven, Dianne Tice · 1998 · Contested

**The starting claim**

Resisting temptation reduces performance on a subsequent self-control task by consuming a limited resource.

**What the research shows**

The radishes-and-cookies experiment made the idea of tiring willpower memorable. Later studies found much smaller effects and different results across protocols; a single reservoir of willpower remains a disputed explanation.

**Evidence that changes the reading**

Hagger and colleagues, 2016: 23 laboratories, 2,141 participants, estimated d = 0.04 with a confidence interval including zero. Dang and colleagues, 2021, found a small effect with a different protocol.

**What you can take away**

Your experience of fatigue deserves to be taken seriously. These findings question one explanation for it and the ability of a short task to reproduce it.

- [Hagger et al. (2016), A Multilab Preregistered Replication](https://doi.org/10.1177/1745691616652873)
- [Dang et al., A Multilab Replication of the Ego Depletion Effect](https://pmc.ncbi.nlm.nih.gov/articles/PMC8186735/)

### The Marshmallow Test

Walter Mischel, Ebbe B. Ebbesen, Antonette Zeiss · 1972 · Supported with limits

**The starting claim**

How long a child waits for a larger reward is associated with later academic achievement.

**What the research shows**

The association exists but shrinks after accounting for family background and early cognitive skills. The test neither determines a child’s future nor shows that waiting causes success.

**Evidence that changes the reading**

Watts, Duncan and Quan, 2018, examined outcomes at age 15, not adult earnings. Among children whose mothers had not completed college, statistical adjustments reduced the initial association by about two thirds.

**What you can take away**

A marshmallow can start a conversation about patience; it cannot deliver a verdict on a child’s character or future.

- [Watts et al. (2018), Revisiting the Marshmallow Test](https://pmc.ncbi.nlm.nih.gov/articles/PMC6050075/)

### Power Posing

Dana Carney, Amy Cuddy, Andy Yap · 2010 · Supported with limits

**The starting claim**

Two minutes in an expansive posture would change hormones and willingness to take risks.

**What the research shows**

The larger replication did not confirm the hormonal promise. Participants did report feeling more powerful, a different outcome from hormonal change or interview success.

**Evidence that changes the reading**

Ranehill and colleagues, 2015, tested 200 people. They found no evidence of changes in testosterone, cortisol or risk-taking on their task.

**What you can take away**

You may like a comfortable posture without it delivering the advertised biology. Feeling confident and performing better need separate measurements.

- [Ranehill et al. (2015), Assessing the Robustness of Power Posing](https://doi.org/10.1177/0956797614553946)
- [University of Zurich, study report](https://www.zne.uzh.ch/en/news/Ranehill-Power-Poses.html)

### Growth Mindset Interventions

Carol Dweck, Lisa Blackwell, Kali Trzesniewski · 2007 · Supported with limits

**The starting claim**

A brief intervention about developing intellectual abilities can improve school achievement.

**What the research shows**

The US national experiment found modest benefits for initially lower-achieving students. Effects also depended on the school environment; a message about effort cannot replace the conditions needed for learning.

**Evidence that changes the reading**

Yeager and colleagues, Nature, 2019: the intervention improved grades among lower-achieving students and increased advanced-mathematics enrolment across the sample.

**What you can take away**

Encouragement can help. When a student struggles, this study does not justify blaming their attitude.

- [Yeager et al. (2019), A national experiment reveals where a growth mindset improves achievement](https://www.nature.com/articles/s41586-019-1466-y)

### The 10,000-Hour Rule

K. Anders Ericsson, Ralf Krampe, Clemens Tesch-Römer · 1993 · Unsupported as stated

**The starting claim**

A universal threshold of 10,000 hours of practice would guarantee expertise.

**What the research shows**

Research on deliberate practice does not establish a threshold guaranteeing success. The relationship with performance varies by domain, and an association alone does not measure the causal benefit of each hour.

**Evidence that changes the reading**

Macnamara, Hambrick and Oswald, 2014, synthesized research across domains. Their 2018 correction changed the overall variance explained from 12% to 14%.

**What you can take away**

Practise purposefully and track improvement. An hour count alone cannot tell you when you will become an expert.

- [Macnamara et al. (2014), Deliberate Practice and Performance](https://doi.org/10.1177/0956797614535810)
- [Corrigendum (2018)](https://doi.org/10.1177/0956797618769891)

### Long-term perseverance: grit

Angela Duckworth, Christopher Peterson, Michael Matthews, Dennis Kelly · 2007 · Supported with limits

**The starting claim**

Perseverance and consistency of long-term interests form a distinct trait called grit that predicts success.

**What the research shows**

Grit overlaps strongly with conscientiousness. Perseverance relates more clearly to performance than consistency of interests and can add predictive information beyond conscientiousness.

**Evidence that changes the reading**

Credé, Tynan and Harms, 2017, synthesized 88 independent samples involving 66,807 people. The findings question the usefulness of a single score combining the two components.

**What you can take away**

Keeping at a task can matter, but a perseverance questionnaire cannot diagnose why someone succeeds or struggles.

- [Credé et al. (2017), Much Ado About Grit](https://doi.org/10.1037/pspp0000102)

### Glucose Model of Self-Control

Matthew Gailliot, Roy Baumeister et al. · 2007 · Unsupported as stated

**The starting claim**

Exerting self-control would lower blood glucose, and consuming sugar would restore the depleted resource.

**What the research shows**

The central predictions of this model were not supported by the research synthesis. Glucose mouth-rinsing findings also offer no secure replacement explanation: the author identified signs of publication bias.

**Evidence that changes the reading**

Dang, Appetite, 2016, separately evaluated glucose depletion, its relationship with later performance and glucose ingestion. None of the three predictions was supported.

**What you can take away**

These experiments do not provide a basis for recommending sugar as a remedy for low willpower.

- [Dang (2016), Testing the role of glucose in self-control](https://pubmed.ncbi.nlm.nih.gov/27492453/)

### How the presentation of choices influences us

Richard Thaler, Cass Sunstein · 2008 · Supported with limits

**The starting claim**

Changing how options are presented can influence choices without removing alternatives.

**What the research shows**

These interventions can work, but results vary with the target and implementation. A small average can matter at scale; it does not promise to solve every social problem.

**Evidence that changes the reading**

DellaVigna and Linos, 2022, analysed 126 trials from two US units covering 23 million people. Average take-up increased by 1.4 percentage points, compared with 8.7 in the academic-paper sample used for comparison.

**What you can take away**

Ask which behaviour changed, by how much and at what cost. This set of trials does not supply a constant for every intervention.

- [DellaVigna & Linos (2022), RCTs to Scale](https://doi.org/10.3982/ECTA18709)

### The Lady Macbeth Effect

Chen-Bo Zhong, Katie Liljenquist · 2006 · Unsupported as stated

**The starting claim**

Recalling an immoral act would make cleaning products more attractive.

**What the research shows**

The link between physical and moral cleanliness is a familiar metaphor. Direct replications of the product-rating task did not confirm the originally reported effect.

**Evidence that changes the reading**

Earp, Everett, Madva and Hamlin, 2014, repeated Zhong and Liljenquist’s Study 2 in three experiments using the original materials. They did not detect the effect in these tests.

**What you can take away**

A compelling metaphor can suggest a hypothesis. Each claim about soap changing moral judgment needs its own test.

- [Earp et al. (2014), Out, damned spot](https://doi.org/10.1080/01973533.2013.856792)
- [Author manuscript](https://www.jimaceverett.com/files/earp2014basp.pdf)

### Cognitive Dissonance (1$ vs 20$ Study)

Leon Festinger, James M. Carlsmith · 1959 · Contested

**The starting claim**

Insufficient external justification for advocating a view against one’s beliefs would encourage attitude change.

**What the research shows**

In the famous experiment, participants received $1 or $20 to describe a dull task as interesting. Later research tests the explanatory mechanism without reducing the whole theory to that single study.

**Evidence that changes the reading**

Vaidis and colleagues, 2024, tested an essay protocol in 39 laboratories. Greater freedom of choice did not produce the predicted attitude difference, although writing a counterattitudinal essay differed from the neutral condition.

**What you can take away**

An attitude change does not automatically identify why it happened. The alternative explanations an experiment rules out matter too.

- [Festinger & Carlsmith (1959), Cognitive consequences of forced compliance](https://doi.org/10.1037/h0041593)
- [Vaidis et al. (2024), A Multilab Replication](https://doi.org/10.1177/25152459231213375)

### Stanford Prison Experiment

Philip Zimbardo · 1971 · Unsupported as stated

**The starting claim**

Merely assigning ordinary people to guard roles would spontaneously turn them into abusers.

**What the research shows**

Abuse occurred in the simulated Stanford prison. Researchers’ instructions, their involvement in events and differences between guards prevent attributing the behaviour solely to the assigned role.

**Evidence that changes the reading**

Le Texier, 2019, analysed archives and participant interviews. He documented researcher interventions and data-collection problems that undermine the account of spontaneous transformation.

**What you can take away**

The case remains important for research ethics and the exercise of power. It cannot establish what anyone would do in a uniform.

- [Le Texier (2019), Debunking the Stanford Prison Experiment](https://pubmed.ncbi.nlm.nih.gov/31380664/)

### Kitty Genovese / 38 Witnesses Myth

A.M. Rosenthal (New York Times) / Bibb Latané, John Darley (1968) · 1964 · Unsupported as stated

**The starting claim**

In 1964, 38 neighbours supposedly watched Kitty Genovese’s murder without anyone trying to help.

**What the research shows**

That account is not supported by the records examined. Witness numbers, what they could see and their interventions were simplified into a story of collective indifference.

**Evidence that changes the reading**

Manning, Levine and Collins, 2007, re-examined the 38-witness account. Their analysis is a historical correction, not an experimental replication of the bystander effect.

**What you can take away**

An inaccurate historical story needs correcting separately from the theory it inspired. Research on helping needs its own assessment.

- [Manning et al. (2007), The Kitty Genovese murder and the social psychology of helping](https://doi.org/10.1037/0003-066X.62.6.555)

### Authority and refusal: Milgram’s experiments

Stanley Milgram · 1963 · Supported with limits

**The starting claim**

In an experiment backed by scientific authority, many participants continue administering what they believe are electric shocks despite the other person’s protests.

**What the research shows**

The pressure to continue is documented, and refusal also occurs. The experiments do not measure the proportion of people willing to kill; the shocks were simulated and the situation depended on the protocol.

**Evidence that changes the reading**

Burger, 2009, conducted a partial replication with 70 adults, stopping at the 150-volt threshold. He observed substantial willingness to continue without testing behaviour through 450 volts.

**What you can take away**

The result invites scrutiny of institutional pressures. It cannot predict an individual’s conduct or remove individual responsibility.

- [Burger (2009), Replicating Milgram](https://pubmed.ncbi.nlm.nih.gov/19209958/)

### Robbers Cave Intergroup Conflict

Muzafer Sherif et al. · 1954 · Supported with limits

**The starting claim**

Competition between groups can foster hostility, while pursuing shared goals can reduce conflict.

**What the research shows**

The camp included researcher-organized competitions followed by problems the boys had to solve together. Conflict did not arise in an untouched setting, and cooperation was built gradually.

**Evidence that changes the reading**

Sherif and colleagues’ 1961 report of the 1954 study describes activities including restoring the water supply. It is the original study report, not an independent replication.

**What you can take away**

Interdependence offers a useful idea for cooperation. One camp study does not establish a universal remedy for historical conflicts.

- [Sherif et al. (1954/1961), Intergroup Conflict and Cooperation, Chapter 7](https://psychclassics.yorku.ca/Sherif/chap7.htm)
- [Sherif et al., conclusions](https://www.yorku.ca/pclassic/Sherif/chap8.htm)

### Asch Conformity Paradigm

Solomon Asch · 1951 · Supported

**The starting claim**

A unanimous majority giving wrong answers can influence someone’s response even on a simple visual task.

**What the research shows**

Asch’s task involved comparing line lengths. Conformity appears in subsequent research, but its strength depends on the situation and cultural context. A public answer alone does not establish changed perception.

**Evidence that changes the reading**

Bond and Smith, 1996, synthesized studies using Asch’s task and found variation across periods and societies. The analysis supports the phenomenon without establishing a universal conformity rate.

**What you can take away**

When a group appears to agree, it is worth asking how answers were collected and whether people could respond independently.

- [Bond & Smith (1996), Culture and conformity](https://doi.org/10.1037/0033-2909.119.1.111)

### From Jerusalem to Jericho (Darley & Batson)

John Darley, Daniel Batson · 1973 · Supported with limits

**The starting claim**

Time pressure can reduce the likelihood of stopping to help someone.

**What the research shows**

Theology students encountered a slumped person on the way to give a talk. Those in a hurry helped less often. The talk’s topic, including the Good Samaritan parable, did not produce the expected difference in this sample.

**Evidence that changes the reading**

Darley and Batson, 1973, compared hurry conditions and speech topics. The cited source is the original experiment; it does not substantiate thousands of replication participants.

**What you can take away**

Circumstances can prevent a good intention from becoming action. The study does not show that moral or religious values have no influence.

- [Darley & Batson (1973), From Jerusalem to Jericho](https://sparq.stanford.edu/sites/g/files/sbiybj19021/files/media/file/darley_batson_1973_-_from_jerusalem_to_jericho.pdf)

### Deindividuation and Cloak of Anonymity

Philip Zimbardo · 1969 · Supported with limits

**The starting claim**

Anonymity and immersion in a group would weaken self-control and encourage norm-breaking.

**What the research shows**

Anonymity has no fixed moral direction. Behaviour also depends on the group’s relevant norms, which can encourage helping or aggression.

**Evidence that changes the reading**

Postmes and Spears, 1998, analysed 60 studies. Results better supported conformity to situation-specific norms than a general loss of norms.

**What you can take away**

To understand an anonymous group, examine which behaviours it encourages.

- [Postmes & Spears (1998), Deindividuation and Antinormative Behavior](https://doi.org/10.1037/0033-2909.123.3.238)

### Who intervenes in a public conflict?

Richard Philpot, Lieke Offermans, Marie Rosenkrantz Lindegaard et al. · 2019 · Supported

**The starting claim**

In observed public conflicts, more bystanders are associated with a greater chance that at least someone intervenes.

**What the research shows**

The chance that a particular person helps and the chance that a victim receives help from someone are different questions. A group can offer a higher chance of intervention even when each member hesitates.

**Evidence that changes the reading**

Philpot and colleagues examined 219 recordings from three countries. At least one bystander intervened in 90.9% of the filmed conflicts. This observation alone does not establish causation or represent every emergency.

**What you can take away**

The figure describes these conflicts, not a personal guarantee. It shows why the perspective used to measure helping matters.

- [Philpot et al., Would I Be Helped?](https://doi.org/10.1037/amp0000469)
- [Published manuscript](https://research.vu.nl/ws/files/259918031/Would_I_Be_Helped_Cross_National_CCTV_Footage_Shows_That_Intervention_Is_the_Norm_in_Public_Conflicts.pdf)

### Minimal Group Paradigm (Henri Tajfel)

Henri Tajfel, Michael Billig, R. P. Bundy, Claude Flament · 1971 · Supported

**The starting claim**

Grouping people by apparently trivial criteria can favour allocating rewards to members of their own group.

**What the research shows**

In minimal-group studies, participants distributed rewards between people identified by group membership. Favouritism could arise without prior conflict; this does not establish that every classification produces hatred.

**Evidence that changes the reading**

Tajfel, Billig, Bundy and Flament, 1971, reported favouritism in reward allocation. The source documents behaviour in these tasks, rather than explaining real-world discrimination on its own.

**What you can take away**

A group label can influence a decision. The size and consequences of that influence need studying in the situation we care about.

- [Tajfel et al. (1971), Social categorization and intergroup behaviour](https://doi.org/10.1002/ejsp.2420010202)

### Individual effort in teamwork

Maximilien Ringelmann / Bibb Latané (1979) · 1913 · Supported

**The starting claim**

People can exert less effort when their contribution is pooled into a collective result.

**What the research shows**

Reduced effort in groups is documented, but depends on task meaning, evaluation of contributions and expectations of colleagues. Coordination losses are a separate explanation for poor group performance.

**Evidence that changes the reading**

Karau and Williams, 1993, synthesized 78 studies and identified many factors that change the effect. The article appeared in the Journal of Personality and Social Psychology.

**What you can take away**

A weak collective result does not prove that team members are lazy. Check whether each understands their contribution and whether the work can be coordinated.

- [Karau & Williams (1993), Social loafing: A meta-analytic review](https://doi.org/10.1037/0022-3514.65.4.681)

### Left-Brain vs Right-Brain Personalities

Distorsionare comercială a lucrărilor lui Roger Sperry · 1981 · Unsupported as stated

**The starting claim**

People would divide into logical, left-brain-dominant types and creative, right-brain-dominant types.

**What the research shows**

Some functions and connections are lateralized. This local specialization does not justify dividing personalities by a dominant hemisphere across the whole brain.

**Evidence that changes the reading**

Nielsen and colleagues, 2013, analysed resting-state images from 1,011 people. They observed local lateralization without evidence for the global left-brained or right-brained types proposed by the myth.

**What you can take away**

You can have both logical and creative abilities. A hemisphere personality quiz does not measure how your brain regions work together.

- [Nielsen et al. (2013), An Evaluation of the Left-Brain vs. Right-Brain Hypothesis](https://doi.org/10.1371/journal.pone.0071275)

### VAK Learning Styles Myth

Walter Barbe, Raymond Swassing, Michael Grinder · 1979 · Unsupported as stated

**The starting claim**

Students would learn better if teaching matched a fixed sensory style, such as visual or auditory.

**What the research shows**

Presentation preferences exist, but do not establish a benefit from matching. The hypothesis requires tests showing that different methods benefit the supposed learner categories differently.

**Evidence that changes the reading**

Pashler, McDaniel, Rohrer and Bjork, 2008, found no adequate basis for using style assessments in general educational practice. This concerns style matching, not every form of teaching adaptation.

**What you can take away**

Choose representations suited to the material and the learner’s actual needs. A style label should not restrict the methods available to them.

- [Pashler et al. (2008), Learning Styles: Concepts and Evidence](https://doi.org/10.1111/j.1539-6053.2009.01038.x)

### The Mozart Effect

Frances Rauscher, Gordon Shaw, Katherine Ky · 1993 · Supported with limits

**The starting claim**

Listening to a Mozart sonata can temporarily improve performance on some spatial tasks.

**What the research shows**

The original finding concerned students and spatial tasks, not children’s intellectual development. Similar advantages occur after other music, and publication bias complicates estimation.

**Evidence that changes the reading**

Pietschnig, Voracek and Formann, 2010, found little evidence for an advantage specific to Mozart. Comparing music with silence asks a different question from comparing it with another piece.

**What you can take away**

Music can remain a pleasure without a promise of higher IQ. An immediate test result does not demonstrate lasting change in intelligence.

- [Pietschnig et al. (2010), Mozart effect–Shmozart effect](https://doi.org/10.1016/j.intell.2010.03.001)

### Can facial expressions change feelings?

Fritz Strack, Leonard Martin, Sabine Stepper · 1988 · Supported with limits

**The starting claim**

Changing facial expression can influence felt emotion, including when a pen held between the teeth mimics a smile.

**What the research shows**

Voluntary expressions and mimicking a smile received support in recent research. Evidence for the pen-in-mouth procedure is less conclusive. A fragile procedure does not invalidate every version of the hypothesis.

**Evidence that changes the reading**

Many Smiles, 2022, tested 3,878 participants across 19 countries with a preregistered plan. Findings differed across methods of producing a smile.

**What you can take away**

This laboratory effect does not justify pressuring someone to smile when distressed and does not establish a treatment for depression.

- [Coles et al. (2022), Many Smiles Collaboration](https://www.nature.com/articles/s41562-022-01458-9)

### 10% Brain Usage Myth

Atribuit eronat lui William James · 1907 · Unsupported as stated

**The starting claim**

We use only 10% of the brain and can unlock the remainder through training.

**What the research shows**

Brain activity continues at rest. Regions contribute differently across tasks; a scan highlighting a few areas often shows a contrast between conditions, leaving ongoing activity out of the picture.

**Evidence that changes the reading**

Raichle and colleagues, 2001, measured baseline brain activity and task-related changes. The study identifies no unused 90% reserve.

**What you can take away**

A brain scan is not a gauge of the percentage of intelligence you are using.

- [Raichle et al. (2001), A default mode of brain function](https://doi.org/10.1073/pnas.98.2.676)

### The Triune Brain (Lizard Brain Myth)

Paul D. MacLean · 1960 · Unsupported as stated

**The starting claim**

Instinct, emotion and reason occupy three brains layered by evolution.

**What the research shows**

The model separates functions too rigidly and misrepresents evolution. Reptiles and mammals share ancestral structures; change involved cell diversification and circuit reorganization.

**Evidence that changes the reading**

Tosches and colleagues, 2018, compared gene expression in reptilian brain cells, identifying evolutionary continuities and differentiated cell types.

**What you can take away**

The lizard-brain metaphor is memorable, but it does not anatomically explain an impulsive purchase.

- [Tosches et al. (2018), Evolution of pallium, hippocampus, and cortical cell types](https://doi.org/10.1126/science.aar4237)

### Priming with words about old age

John Bargh, Mark Chen, Lara Burrows · 1996 · Unsupported as stated

**The starting claim**

Unscrambling sentences containing words associated with old age makes participants walk more slowly.

**What the research shows**

The slower-walking effect was not convincingly reproduced with automated timing. Experimenter expectations affected results in a subsequent test. This does not refute every form of priming.

**Evidence that changes the reading**

Doyen and colleagues, 2012, used sensors to time walking and separately manipulated experimenter expectations.

**What you can take away**

A detail such as who holds the stopwatch can change the interpretation of a striking experiment.

- [Doyen et al. (2012), Behavioral priming: It’s all in the mind, but whose mind?](https://doi.org/10.1371/journal.pone.0029081)

### Mirror Neurons Overreach

Giacomo Rizzolatti, Vittorio Gallese et al. · 1996 · Supported with limits

**The starting claim**

Cells responding both to performing and observing an action could explain human empathy.

**What the research shows**

Cells with these responses exist. Finding them does not establish that they alone explain empathy, language or understanding intentions. Their role requires testing for each function.

**Evidence that changes the reading**

Mukamel and colleagues, 2010, recorded 1,177 neurons in 21 patients observing and performing gestures. Some cells responded in both situations.

**What you can take away**

A cell responding when you see a gesture is not, by itself, an explanation of compassion.

- [Mukamel et al. (2010), Single-neuron responses in humans during execution and observation of actions](https://doi.org/10.1016/j.cub.2010.02.045)

### Memories of shocking news

Roger Brown, James Kulik · 1977 · Supported with limits

**The starting claim**

The moment you hear shocking news is preserved with photographic precision.

**What the research shows**

Memories can remain vivid and convincing even as reported details change. Consistency across retellings is itself different from accuracy against the event.

**Evidence that changes the reading**

Talarico and Rubin, 2003, followed 54 students after September 11. Consistency declined for both the news and an everyday event; confidence held up better for the former.

**What you can take away**

A person’s sincerity and certainty deserve respect without guaranteeing every remembered detail.

- [Talarico & Rubin (2003), Confidence, not consistency, characterizes flashbulb memories](https://doi.org/10.1111/1467-9280.02453)

### Repressed and recovered memories

Mișcarea clinică de hipnoză și regresie terapeutică · 1980 · Unsupported as stated

**The starting claim**

Hypnosis and suggestive questioning can reliably recover trauma completely blocked in the unconscious.

**What the research shows**

Suggestion can contribute to false memories. A memory recalled after a long interval can nevertheless have independent corroboration; its timing neither proves nor disproves it.

**Evidence that changes the reading**

Loftus and Pickrell, 1995, studied a suggested memory of getting lost in a shopping centre. Geraerts and colleagues, 2007, found corroboration for some spontaneously recovered abuse memories.

**What you can take away**

Listening without leading questions protects both the person and the quality of the evidence.

- [Loftus & Pickrell (1995), The formation of false memories](https://doi.org/10.3928/0048-5713-19951201-07)
- [Geraerts et al. (2007), The reality of recovered memories](https://doi.org/10.1111/j.1467-9280.2007.01940.x)

### Broken Windows Theory

James Q. Wilson, George L. Kelling · 1982 · Supported with limits

**The starting claim**

Addressing minor disorder prevents serious crime.

**What the research shows**

The intervention matters. Programs solving local problems with communities show more favorable results than aggressive enforcement against minor offenses. These evaluations do not by themselves test the full theoretical mechanism.

**Evidence that changes the reading**

The Braga, Welsh and Schnell review, 2019, finds crime reductions associated with community interventions, without a significant reduction for aggressive order maintenance.

**What you can take away**

Repairing a place and penalizing its users are different interventions with different evidence.

- [Braga et al. (2019), Disorder policing to reduce crime: A systematic review](https://doi.org/10.1002/cl2.1050)

### Stereotype Threat Overreach

Claude Steele, Joshua Aronson · 1995 · Contested

**The starting claim**

Fear of confirming a negative stereotype can lower test performance.

**What the research shows**

There are findings supporting the hypothesis, but effect size and generalizability are disputed. An analysis about girls and mathematics cannot settle the effect for every group, examination or form of discrimination.

**Evidence that changes the reading**

Flore and Wicherts, 2015, analyzed 47 effect sizes among girls under 18. The average was small, with publication-bias indicators casting doubt on its estimate.

**What you can take away**

Uncertainty about a particular intervention does not make the experience of discrimination imaginary.

- [Flore & Wicherts (2015), Does stereotype threat influence performance of girls in stereotyped domains?](https://doi.org/10.1016/j.jsp.2014.10.002)

### Implicit Association Test (IAT)

Anthony Greenwald, Debbie McGhee, Jordan Schwartz · 1998 · Supported with limits

**The starting claim**

Differences in category-association speed could predict an individual’s discriminatory behavior.

**What the research shows**

The IAT measures relative associations under test conditions. Its relationships with behavior vary and do not justify a confident individual verdict about character or future actions.

**Evidence that changes the reading**

Oswald and colleagues, 2013, find poor prediction. Kurdi and colleagues, 2019, find significant but heterogeneous associations in a broader synthesis. Conclusions depend on criteria and methods.

**What you can take away**

A score can open a conversation about associations; it should not become a moral label.

- [Oswald et al. (2013), Predicting ethnic and racial discrimination](https://doi.org/10.1037/a0032734)
- [Kurdi et al. (2019), Relationship between the IAT and intergroup behavior](https://doi.org/10.1037/amp0000364)

### Income inequality and mental health

Richard Wilkinson, Kate Pickett (The Spirit Level) · 2009 · Supported with limits

**The starting claim**

Income inequality generally harms mental health beyond personal income.

**What the research shows**

The recent synthesis does not support a universal negative relationship. Associations vary with context, including income and inflation; the result does not refute the health effects of poverty.

**Evidence that changes the reading**

Sommet and colleagues, published online in 2025, combine 168 studies. The average mental-health association is near zero after publication-bias correction, but remains adverse in low-income samples.

**What you can take away**

Inequality between people and a person’s lack of money require distinct questions and measurements.

- [Sommet et al. (2025), No meta-analytical effect of economic inequality on well-being or mental health](https://doi.org/10.1038/s41586-025-09797-z)

### Scarcity Brain Drain Effect

Sendhil Mullainathan, Eldar Shafir · 2013 · Supported with limits

**The starting claim**

Financial concerns reduce attention available for other tasks.

**What the research shows**

Initial research found lower performance under certain conditions. Later results are uneven; the popular 13-IQ-point figure is not a fixed loss of intelligence caused by poverty.

**Evidence that changes the reading**

Mani and colleagues, 2013, studied financial scenarios and farmers before and after harvest. Carvalho, Meier and Wang, 2016, found no cognitive differences before versus after payday.

**What you can take away**

Financial difficulty is not a diagnosis of someone’s intellectual capacity.

- [Mani et al. (2013), Poverty impedes cognitive function](https://doi.org/10.1126/science.1238041)
- [Carvalho et al. (2016), Poverty and economic decision-making](https://doi.org/10.1257/aer.20140481)

### Subliminal Advertising Panic

James Vicary · 1957 · Unsupported as stated

**The starting claim**

Messages hidden in a film automatically increase drink and popcorn sales.

**What the research shows**

Vicary’s story provides no credible evidence for the advertised sales increases. Laboratory research on subliminal stimuli is a separate question, with condition-dependent effects.

**Evidence that changes the reading**

Karremans, Stroebe and Claus, 2006, reported an effect of a subliminal brand name on drink choice among thirsty participants.

**What you can take away**

A limited influence on a laboratory choice does not demonstrate covert control over everyday shopping.

- [Karremans et al. (2006), Beyond Vicary’s fantasies](https://doi.org/10.1016/j.jesp.2005.12.002)

### The Paradox of Choice (Jam Study)

Sheena Iyengar, Mark Lepper · 2000 · Supported with limits

**The starting claim**

A large assortment can discourage buying, as the comparison of 24 versus 6 jams suggested.

**What the research shows**

More options do not automatically cause paralysis. Studies find different results; a near-zero average can combine situations where variety helps and where it hinders.

**Evidence that changes the reading**

Scheibehenne, Greifeneder and Todd, 2010, synthesized 50 experiments with 5,036 participants. The average effect was near zero, with substantial variation between studies.

**What you can take away**

For a difficult choice, clear criteria may matter more than arbitrarily shortening the list.

- [Iyengar & Lepper (2000), When choice is demotivating](https://doi.org/10.1037/0022-3514.79.6.995)
- [Scheibehenne et al. (2010), Can there ever be too many options?](https://doi.org/10.1086/651235)

### Intergroup Contact Hypothesis

Gordon Allport · 1954 · Supported with limits

**The starting claim**

Contact between members of different groups can reduce prejudice.

**What the research shows**

Evidence supports an average reduction in prejudice, with variation across interventions. The favorable conditions proposed by the theory are not requirements proven necessary in every case.

**Evidence that changes the reading**

Paluck, Green and Green, 2019, found 27 randomized studies with delayed outcomes. Effects on racial or ethnic prejudice were weaker, with gaps in evidence about adults.

**What you can take away**

It is worth testing which kind of encounter helps, whom it helps and for how long.

- [Paluck et al. (2019), The contact hypothesis re-evaluated](https://doi.org/10.1017/bpp.2018.25)

### The Dunning–Kruger effect

Justin Kruger, David Dunning · 1999 · Contested

**The starting claim**

Poor performers overestimate themselves because they also lack the skills to recognize their mistakes.

**What the research shows**

The classic self-assessment pattern can be amplified by measurement error and regression to the mean. The contribution of a specific metacognitive mechanism remains disputed; popular graphs about experts and fools exceed the evidence.

**Evidence that changes the reading**

Gignac and Zajenkowski, 2020, tested measured and self-assessed intelligence in 929 people without finding the nonlinear pattern predicted by the classic explanation.

**What you can take away**

Compare your estimate with a checkable result. Labeling an opponent does not test your own judgment.

- [Gignac & Zajenkowski (2020), The Dunning–Kruger effect is mostly a statistical artefact](https://doi.org/10.1016/j.intell.2020.101449)

### Big Five Personality Structure

Lewis Goldberg, Paul Costa, Robert McCrae · 1990 · Supported with limits

**The starting claim**

Five dimensions describe much of the variation in personality.

**What the research shows**

The model is useful and well studied, particularly in literate populations. Its dimensions are continuous, and their structure and measurement do not automatically transfer to every culture.

**Evidence that changes the reading**

Gurven and colleagues, 2013, did not reproduce the standard structure in responses from Bolivia’s Tsimane population. That limitation matters when presenting an instrument as universal.

**What you can take away**

A profile describes tendencies; it neither determines your future nor confines you to a personality type.

- [Gurven et al. (2013), How universal is the Big Five?](https://doi.org/10.1037/a0030841)

### Spaced repetition and retrieval practice

Hermann Ebbinghaus / Roediger & Karpicke (2006) · 1885 · Supported

**The starting claim**

Spacing study sessions and trying to retrieve material from memory can improve long-term retention.

**What the research shows**

These are two distinct practices with supporting evidence. Benefits depend on material, timing and assessment; rereading can feel easier and help immediately without offering the same later advantage.

**Evidence that changes the reading**

Roediger and Karpicke, 2006, found a testing advantage on delayed assessments. Cepeda and colleagues, 2006, synthesized evidence on distributed study.

**What you can take away**

Close the material, try to explain what you remember, check your answer and return after an interval.

- [Roediger & Karpicke (2006), Test-enhanced learning](https://doi.org/10.1111/j.1467-9280.2006.01693.x)
- [Cepeda et al. (2006), Distributed practice in verbal recall tasks](https://doi.org/10.1037/0033-2909.132.3.354)

### How much fits in working memory?

George A. Miller / Nelson Cowan (2001) · 1956 · Supported with limits

**The starting claim**

Under certain conditions, short-term memory holds approximately four information units.

**What the research shows**

Capacity is limited, but the number depends on the task, grouping and prior knowledge. Four is a useful estimate under controlled conditions, not a universal ceiling on thought.

**Evidence that changes the reading**

Cowan, 2001, reexamines research and proposes about four units when rehearsal and long-term-memory support are restricted.

**What you can take away**

Write down task steps and group information meaningfully. An external list reduces the burden of holding everything in mind.

- [Cowan (2001), The magical number 4 in short-term memory](https://doi.org/10.1017/S0140525X01003922)

### The Stroop effect

John Ridley Stroop · 1935 · Supported

**The starting claim**

Naming ink color becomes harder when the printed word names a different color.

**What the research shows**

In fluent readers, word information can interfere with the required response. The difference depends on the task version; the test alone neither localizes a brain problem nor makes a diagnosis.

**Evidence that changes the reading**

Stroop, 1935, compared reading and color naming with and without conflicting verbal information. Incompatible words slowed color naming.

**What you can take away**

Hesitating over RED printed in blue illustrates ordinary competition between responses.

- [Stroop (1935), Studies of interference in serial verbal reactions](https://doi.org/10.1037/h0054651)

### Cognitive Behavioral Therapy Efficacy

Aaron Beck, A. John Rush et al. · 1977 · Supported

**The starting claim**

Cognitive behavioral therapy can reduce depression symptoms compared with control conditions.

**What the research shows**

Randomized trials support efficacy. Benefits vary between people and comparisons; superiority over other psychotherapies is unclear. Treatment includes behavioral practices as well as work on thoughts.

**Evidence that changes the reading**

Cuijpers and colleagues, 2023, synthesize 409 trials with 52,702 patients. They find benefits against conditions including waitlist and usual care.

**What you can take away**

If a therapy does not help enough, the result deserves discussion and treatment adjustment; it is not a personal failure.

- [Cuijpers et al. (2023), Cognitive behavior therapy versus control conditions and other treatments for depression](https://doi.org/10.1002/wps.21069)

### The pull of negative information

Roy Baumeister, Ellen Bratslavsky, Catrin Finkenauer, Kathleen Vohs · 2001 · Supported with limits

**The starting claim**

Negative information can attract more attention than positive information.

**What the research shows**

The asymmetry appears in studied contexts without being identical across stimuli, people or memory. A click effect establishes neither an evolutionary explanation nor how every algorithm works.

**Evidence that changes the reading**

Robertson and colleagues, 2023, analyzed randomized Upworthy headline tests. For an average-length headline, an additional negative word increased click-through rate by 2.3% in relative terms.

**What you can take away**

A headline attracting you through fear is not necessarily more relevant or better supported.

- [Robertson et al. (2023), Negativity drives online news consumption](https://doi.org/10.1038/s41562-023-01538-4)

### Heritability of Psychological Traits (Twin Studies)

Thomas Bouchard, David Lykken, Matthew McGue et al. · 1990 · Supported

**The starting claim**

Genetic differences contribute to variation in many human traits.

**What the research shows**

Twin studies support genetic contributions, with estimates depending on population, environment and model assumptions. Heritability does not say how much of a person is genetic or mean that a trait cannot change.

**Evidence that changes the reading**

Polderman and colleagues, 2015, synthesize 2,748 publications on human traits, many nonpsychological. Twin pairs in the total partly overlap across measurements.

**What you can take away**

An estimate about differences within a population cannot determine a child’s future.

- [Polderman et al. (2015), Meta-analysis of the heritability of human traits](https://doi.org/10.1038/ng.3285)

### Anchoring and Adjustment Heuristic

Amos Tversky, Daniel Kahneman · 1974 · Supported

**The starting claim**

A starting value can pull subsequent numerical estimates toward it.

**What the research shows**

The effect occurs in many estimation tasks. In the classic demonstration, a wheel rigged to show 10 or 65 preceded estimates of African countries’ UN share. Influence is not inevitable for every person.

**Evidence that changes the reading**

Klein and colleagues, 2014, found strong anchoring evidence in Many Labs 1, a project with 36 samples. Its protocol did not exactly repeat the wheel demonstration.

**What you can take away**

Before a negotiation, prepare an independent estimate and the criteria behind it.

- [Tversky & Kahneman (1974), Judgment under uncertainty](https://doi.org/10.1126/science.185.4157.1124)
- [Klein et al. (2014), Investigating variation in replicability](https://doi.org/10.1027/1864-9335/a000178)

### Seeking confirmation: the 2–4–6 problem

Peter Cathcart Wason · 1960 · Supported

**The starting claim**

When discovering a rule, people can favor examples matching their first hypothesis and miss alternatives.

**What the research shows**

The task shows a difficulty in hypothesis testing without proving that people always reject contrary evidence. Results also depend on how the problem is constructed.

**Evidence that changes the reading**

Wason, 1960, used the 2–4–6 sequence and a hidden rule, not the four-card task. Only 6 of 29 participants found the rule without an earlier incorrect conclusion.

**What you can take away**

Look for an example that would distinguish your explanation from a plausible alternative.

- [Wason (1960), On the failure to eliminate hypotheses in a conceptual task](https://doi.org/10.1080/17470216008416717)

## The next time a headline explains “how your mind works”

Start with what participants did and what was measured. Then ask whether other researchers found something similar, how large the effect was, and how closely the situation resembles your life. A finding can stay interesting as its promise becomes more modest.

## How we selected and read these cases

This is an editorial selection of 48 familiar experiments, theories, and practices. It is not a representative sample of psychology or a systematic review of the entire literature. We retained the original edition’s subjects and checked their claims against original studies, replications, meta-analyses, and archival research. The year in each case is a historical reference, not the date of its latest evidence.

Labels summarize our reading of the cited sources at the revision date. New research may change them. Sample sizes appear beside identified studies; a meta-analysis and a direct replication are different kinds of evidence. The therapy case describes research evidence; treatment decisions require assessment of a person’s clinical situation with a professional.

### What changed in this edition

We replaced categorical verdicts, corrected references and attributions, and removed participant totals without an identifiable source. The lab now uses independent observations preserved across analyses and a separate check on fresh data. All 48 subjects remain available.

- [Open Science Collaboration (2015). Estimating the reproducibility of psychological science](https://doi.org/10.1126/science.aac4716)
- [Simmons, Nelson & Simonsohn (2011). False-Positive Psychology](https://doi.org/10.1177/0956797611417632)
- [Wasserstein & Lazar (2016). The ASA Statement on p-Values](https://doi.org/10.1080/00031305.2016.1154108)
