Do people learn like Bayesians when rewards suddenly change?
People chose like Bayesian learners who dislike uncertainty, but only when told that reward odds could suddenly jump; otherwise simple trial-and-error learning fit just as well.
Source
Risk, unexpected uncertainty, and estimation uncertainty: Bayesian learning in unstable settings
Study at a glance
- Design
- Human experiment — Model comparison (Bayesian learner with forgetting vs Rescorla-Wagner and Pearce-Hall) on choices from a six-arm restless bandit, plus a new experiment manipulating how much task structure participants were told.
- N
- 43 undergraduates in the no-information treatment; the numbers in the other two treatments are garbled in the extracted text, and treatments 1 and 2 pooled gave 75 cases. Reanalysed data from an earlier experiment are also used.
- Population
- Undergraduates at Ecole Polytechnique Fédérale Lausanne
- Outcome
- Model fit (log-likelihood, BIC) of learning models to trial-by-trial choices; debriefing reports of noticing jumps
Structured fields used in claim comparison tables when every cited study has a complete layer.
What they did
Participants played a board game with six options whose win and loss probabilities changed abruptly without warning, for roughly 500 trials. The authors fitted a Bayesian model that separately tracks risk, estimation uncertainty and the chance of a sudden change, and compared it with model-free reinforcement-learning models. They also tested whether estimation uncertainty acted as a bonus that encourages exploring or a penalty that discourages it, and ran three treatments that varied how much of the task structure participants were told.
What they found
When fully informed about the task, participants' choices were better fitted by the Bayesian model than by reinforcement learning, and the Pearce-Hall model fitted worst. Adding a penalty for estimation uncertainty improved the fit, while an exploration bonus made it worse, so people avoided rather than sought uncertain options. When participants were not told that probabilities could jump, the Bayesian model no longer beat reinforcement learning, and debriefing showed most had not noticed the jumps even though red options jumped four times more often.
The limits
What it doesn't show
Key statistics (exact p-values, some group sizes) are missing from the extracted text, so the size of the model differences cannot be checked here. The sample is undergraduates at a single university, and participants in the first treatment could opt into the others, so treatments were not independent random groups. Better model fit shows the Bayesian account describes choices well, not that the brain literally computes these quantities; the neural claims in the discussion come from other studies, not new data.
Key terms
- Restless bandit
- A choice task with several options whose reward probabilities change over time, so the best option must be relearned.
- Risk
- Uncertainty that remains even when the outcome probabilities are known exactly.
- Estimation uncertainty (ambiguity)
- Uncertainty about what the outcome probabilities are, which shrinks as you gather data.
- Unexpected uncertainty
- The chance that the environment has suddenly changed, meaning old learning should be discarded.
- Model-free reinforcement learning
- Learning option values directly from prediction errors without a model of how outcomes are generated.
- Softmax
- A rule that turns option values into choice probabilities, balancing exploiting the best option and exploring others.
Flashcards
0 of 10 answers reviewed
Research intelligence for this paper
See its role on concept claims, tensions it is part of, placement history, and related discoveries.
Quiz yourself
Which kind of uncertainty signals that past learning should be partly forgotten?
Common questions
Why does it matter to separate three kinds of uncertainty?
Each should change learning differently: risk should be ignored, estimation uncertainty calls for more learning, and a likely sudden change means you should start learning again.
Did people explore uncertain options more?
No. The model fitted better when uncertainty lowered an option's value, consistent with ambiguity aversion rather than an exploration bonus.
Why did uninformed participants look like simple reinforcement learners?
Most did not notice the sudden jumps and blamed bad runs on chance, so they had no model of change to use; the authors suggest people fall back on model-free learning under structural uncertainty.
More on Uncertainty and calibration