Skip to content
PaperFren

Do people learn more from outcomes that confirm their choice?

Open paper intelligence

People learned more from news that their choice was right, whether it came from the option they picked or the one they skipped, and this bias made them slower to adapt when rewards switched.

Source

Confirmation bias in human reinforcement learning: Evidence from counterfactual feedback processing

Palminteri S, Lefebvre G, Kilford EJ, et al. · PLoS computational biology · 2017

doi.org/10.1371/journal.pcbi.1005684Read the full paper ↗133 citationscc by

Study at a glance

Design
Human experiment — Two within-subject probabilistic two-armed bandit experiments (chosen-outcome feedback only vs chosen plus forgone feedback), analysed with fitted reinforcement-learning models and BIC model comparison.
N
N=40 · 20 healthy adults in Experiment 1 and a different 20 in Experiment 2; 192 trials each.
Population
Healthy adult volunteers
Outcome
Fitted learning rates for positive/negative factual and counterfactual prediction errors; model fit; correct and preferred choice rates

Structured fields used in claim comparison tables when every cited study has a complete layer.

What they did

Two groups of 20 adults repeatedly chose between pairs of symbols that won or lost a point with fixed or reversing probabilities. In Experiment 1 they saw only the outcome of their chosen option; in Experiment 2 they also saw what the unchosen option would have paid. The authors fitted reinforcement-learning models with separate learning rates for good and bad surprises from chosen and unchosen options, and compared simpler models by BIC.

What they found

For chosen options, learning rates were higher after better-than-expected outcomes, replicating a known positivity asymmetry. For unchosen options the asymmetry reversed: people learned more when the forgone option did worse than expected, a strong valence-by-type interaction in Experiment 2. A two-parameter model with one rate for confirming and one for disconfirming outcomes fitted best, and participants with bigger biases performed worse after the reversal and stuck more to one option when both paid equally.

The limits

What it doesn't show

Each experiment had only 20 young, healthy participants, and the high/low bias comparisons are median splits on small groups. Learning rates are model-derived estimates, so the conclusion depends on the Rescorla-Wagner model family being a good description of behaviour. The task used simple point rewards over a single session, so it is untested whether the bias holds for real-money stakes, rare large rewards, observational learning or the belief perseverance the authors speculate about. The claim that the bias may be adaptive through self-esteem is speculation, not tested here.

Key terms

Prediction error
The difference between the outcome received and the outcome expected; positive if better than expected, negative if worse.
Learning rate
A model parameter setting how much a single prediction error updates the value estimate of an option.
Counterfactual learning
Updating beliefs using the outcome of an option you did not choose (the forgone outcome).
Confirmation bias (in learning)
Giving more weight to outcomes that support your current choice (good chosen or bad unchosen results) than to outcomes that contradict it.
Bayesian Information Criterion (BIC)
A model-comparison score that rewards goodness of fit but penalises extra parameters.
Rescorla-Wagner model
A simple learning rule where value estimates move towards outcomes in proportion to the prediction error times a learning rate.

Flashcards

1 / 10

0 of 10 answers reviewed

Research intelligence for this paper

See its role on concept claims, tensions it is part of, placement history, and related discoveries.

Open paper intelligence

Quiz yourself

1 / 6

What distinguishes a confirmation bias from a positivity bias in this study?

Common questions

Why was Experiment 2 needed if Experiment 1 already showed a bias?

With only chosen outcomes, a positivity bias and a confirmation bias predict the same thing. Showing forgone outcomes separates them, because they predict opposite asymmetries for unchosen options.

Is the bias harmful?

In this task, yes: higher-bias participants were worse at switching after the reward probabilities reversed. The authors suggest it may still be useful elsewhere, for example by supporting confidence, but that was not tested.

Could the result be an artefact of model fitting?

The authors checked this with parameter recovery on simulated agents, found learning rates were recovered reasonably well, and found no spurious negative correlation between the confirmatory and disconfirmatory rates.

More on Reinforcement learning