Reinforcement learning
Do people learn more from outcomes that confirm their choice?
Open access · cc by · source: Europe PMC
People learned more from news that their choice was right, whether it came from the option they picked or the one they skipped, and this bias made them slower to adapt when rewards switched.
Study at a glance
- Design
- Human experiment — Two within-subject probabilistic two-armed bandit experiments (chosen-outcome feedback only vs chosen plus forgone feedback), analysed with fitted reinforcement-learning models and BIC model comparison.
- N
- N=40 · 20 healthy adults in Experiment 1 and a different 20 in Experiment 2; 192 trials each.
- Population
- Healthy adult volunteers
- Outcome
- Fitted learning rates for positive/negative factual and counterfactual prediction errors; model fit; correct and preferred choice rates
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
For chosen options, learning rates were higher after better-than-expected outcomes, replicating a known positivity asymmetry. For unchosen options the asymmetry reversed: people learned more when the forgone option did worse than expected, a strong valence-by-type interaction in Experiment 2. A two-parameter model with one rate for confirming and one for disconfirming outcomes fitted best, and participants with bigger biases performed worse after the reversal and stuck more to one option when both paid equally.
Methodology
Two groups of 20 adults repeatedly chose between pairs of symbols that won or lost a point with fixed or reversing probabilities. In Experiment 1 they saw only the outcome of their chosen option; in Experiment 2 they also saw what the unchosen option would have paid. The authors fitted reinforcement-learning models with separate learning rates for good and bad surprises from chosen and unchosen options, and compared simpler models by BIC.
Limitations
Each experiment had only 20 young, healthy participants, and the high/low bias comparisons are median splits on small groups. Learning rates are model-derived estimates, so the conclusion depends on the Rescorla-Wagner model family being a good description of behaviour. The task used simple point rewards over a single session, so it is untested whether the bias holds for real-money stakes, rare large rewards, observational learning or the belief perseverance the authors speculate about. The claim that the bias may be adaptive through self-esteem is speculation, not tested here.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
We update in whichever direction flatters our choice.
People learn more from outcomes that confirm their choice: for chosen options, learning rates were higher after better-than-expected outcomes, and for unchosen options people learned more when the forgone option did worse than expected; participants with bigger biases performed worse after reversal.
Evidence for the claim as stated.
Related papers in this topic
Same topic cluster — not a recommendation engine.