Skip to content
PaperFren

Concept

Learning biases in human reinforcement learning

5 studies1 discoveryEvidence last moved Sep 27, 2026

Human reinforcement learning is often modelled with prediction-error rules like Rescorla-Wagner, but people don't update evenly: learning rates differ by outcome valence, action type, social context and age. The evidence here is from small laboratory human experiments with computational model fitting.

These 'biases' are what make human learners unlike textbook algorithms, and model parameters are increasingly used as traits. Knowing that they come from small samples and model-dependent estimates stops students overreading them.

Studies

5

Findings

4

4 supporting · 0 challenging · 0 qualifying citations

Open tensions

1

Latest change

Concept page published

Learning biases in human reinforcement learning

Currently

What we know

  1. We update in whichever direction flatters our choice.
  2. Valence interacts with approach/withdrawal, not just with value.
  3. Adolescents use a simpler learning strategy, not less motivation.
  4. Learning rates are not fixed; they rise when the world seems unstable.

Largest unresolved question

Whether punishment carries a real learning signal: in the approach/withdrawal task punishment's average learning effect was indistinguishable from zero, while adults in the adolescence study used punishment-context information that adolescents lacked.

Common misconceptions

  • A fitted learning rate is a directly measured property of the brain.

    Learning rates are model-derived; the confirmation-bias and hierarchical-Gaussian-filter results depend on which model family was tested, and model selection only picks the best of those tried.

  • These biases are established population facts.

    Samples were 15–46 people (the adviser study all male; adolescence study cross-sectional with 38 people), so they are well-controlled but small lab findings.

Related