Concept · artificial-intelligence
Learning biases in human reinforcement learning
5 studies1 discoveryEvidence last moved Sep 27, 2026
Human reinforcement learning is often modelled with prediction-error rules like Rescorla-Wagner, but people don't update evenly: learning rates differ by outcome valence, action type, social context and age. The evidence here is from small laboratory human experiments with computational model fitting.
These 'biases' are what make human learners unlike textbook algorithms, and model parameters are increasingly used as traits. Knowing that they come from small samples and model-dependent estimates stops students overreading them.
Studies
5
Findings
4
4 supporting · 0 challenging · 0 qualifying citations
Open tensions
1
Latest change
Concept page published
Learning biases in human reinforcement learning
Currently
What we know
- We update in whichever direction flatters our choice.
- Valence interacts with approach/withdrawal, not just with value.
- Adolescents use a simpler learning strategy, not less motivation.
- Learning rates are not fixed; they rise when the world seems unstable.
Largest unresolved question
Whether punishment carries a real learning signal: in the approach/withdrawal task punishment's average learning effect was indistinguishable from zero, while adults in the adolescence study used punishment-context information that adolescents lacked.
Common misconceptions
A fitted learning rate is a directly measured property of the brain.
Learning rates are model-derived; the confirmation-bias and hierarchical-Gaussian-filter results depend on which model family was tested, and model selection only picks the best of those tried.
These biases are established population facts.
Samples were 15–46 people (the adviser study all male; adolescence study cross-sectional with 38 people), so they are well-controlled but small lab findings.
Related
Claim ledger
What the evidence shows
Drawn from 5 studies in this library. Mix labels say which citation roles are present; they are not a strength score. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope.
We update in whichever direction flatters our choice.
People learn more from outcomes that confirm their choice: for chosen options, learning rates were higher after better-than-expected outcomes, and for unchosen options people learned more when the forgone option did worse than expected; participants with bigger biases performed worse after reversal.
Valence interacts with approach/withdrawal, not just with value.
Reward and punishment are not mirror images: reward-predicting cues increased 'go' responses for approach but reduced them for withdrawal (approach effect in 45 of 46 people), and rewards moved learning much more than punishments, whose average effect was indistinguishable from zero.
Adolescents use a simpler learning strategy, not less motivation.
Learning strategies change with development: basic Q-learning fitted adolescents best while adults were best fit by a model including punishment-context and counterfactual modules; adolescents did not benefit from seeing unchosen outcomes, though both groups learned equally in the simplest context.
Learning rates are not fixed; they rise when the world seems unstable.
When learning about others, people adjust learning speed to perceived volatility: a hierarchical Gaussian filter beat Rescorla-Wagner in a social advice task, and its trial-by-trial estimates tracked players' explicit ratings of the adviser.
Debates
Tensions and limits
Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes.
Whether punishment carries a real learning signal: in the approach/withdrawal task punishment's average learning effect was indistinguishable from zero, while adults in the adolescence study used punishment-context information that adolescents lacked.
Evidence against
PaperFren reads this as a limit on how far one study travels — different assays, populations, or outcomes — not a forced fight between papers.
Timeline
How understanding moved
Study years are when the paper was published. Evidence edits are dated changes to this page's claims. Explanations are when PaperFren added a Discovery — not a claim that the science happened that day.
2026
- The standard test of planning versus habit barely rewards planning, and can be passed without it
Concept page published
Learning biases in human reinforcement learning
Change log
What changed
Dated edits to this page's evidence: studies added or removed from a claim, claims added or withdrawn, and new explanations tagged here. Rewordings are not listed.
- Concept page published
Papers
5 studies in this library bear on Learning biases in human reinforcement learning, ordered by citations.
- Do reward cues push us to approach rather than just to act?
Cues that predicted money made people more likely to approach but less likely to withdraw, so Pavlovian cues bias specific kinds of action rather than simply energising behaviour.
- Are habits chunked action sequences run by a goal-directed boss?
People's habitual choices were better explained as whole pre-packaged action sequences chosen by a goal-directed system than as separate single actions valued by a model-free habit system.
- How do people learn whether an adviser is trying to help them?
People's choices were best explained by a learning model that tracks not only how accurate an adviser is but also how quickly the adviser's intentions are changing.
- Do people learn more from outcomes that confirm their choice?
People learned more from news that their choice was right, whether it came from the option they picked or the one they skipped, and this bias made them slower to adapt when rewards switched.
- Do teenagers learn from rewards and punishments like adults?
Adolescents' learning was best explained by a simple reward-tracking algorithm, while adults also learned from outcomes they didn't choose and from context, which helped them avoid punishments.
Compare studies
Select 2–10 studies. Design and N are labels, not a ranking.
Nothing selected yet.
Questions
What is still open
Whether punishment carries a real learning signal: in the approach/withdrawal task punishment's average learning effect was indistinguishable from zero, while adults in the adolescence study used punishment-context information that adolescents lacked.
Ask PaperFren about Learning biases in human reinforcement learning
Study this conceptflashcards and short-answer questions
Describe the evidence for confirmation bias in reinforcement learning and one limitation.
With counterfactual feedback, participants had higher learning rates for positive surprises on chosen options and negative surprises on unchosen ones, the pattern a confirmatory update predicts. A two-rate confirm/disconfirm model fit best, and bigger biases went with worse reversal performance. Each experiment had only 20 young adults and the learning rates depend on the Rescorla-Wagner model family.
How did adolescents' learning differ from adults' in the computational development study?
Model comparison favoured simple Q-learning for adolescents and a fuller model with punishment-context and counterfactual components for adults. Adolescents improved less at avoiding punishment and ignored unchosen outcomes, while both groups did equally well in the simplest condition. The design was cross-sectional with 38 participants.