Uncertainty and calibration
Do people learn like Bayesians when rewards suddenly change?
Open access · cc by · source: Europe PMC
People chose like Bayesian learners who dislike uncertainty, but only when told that reward odds could suddenly jump; otherwise simple trial-and-error learning fit just as well.
Study at a glance
- Design
- Human experiment — Model comparison (Bayesian learner with forgetting vs Rescorla-Wagner and Pearce-Hall) on choices from a six-arm restless bandit, plus a new experiment manipulating how much task structure participants were told.
- N
- 43 undergraduates in the no-information treatment; the numbers in the other two treatments are garbled in the extracted text, and treatments 1 and 2 pooled gave 75 cases. Reanalysed data from an earlier experiment are also used.
- Population
- Undergraduates at Ecole Polytechnique Fédérale Lausanne
- Outcome
- Model fit (log-likelihood, BIC) of learning models to trial-by-trial choices; debriefing reports of noticing jumps
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
When fully informed about the task, participants' choices were better fitted by the Bayesian model than by reinforcement learning, and the Pearce-Hall model fitted worst. Adding a penalty for estimation uncertainty improved the fit, while an exploration bonus made it worse, so people avoided rather than sought uncertain options. When participants were not told that probabilities could jump, the Bayesian model no longer beat reinforcement learning, and debriefing showed most had not noticed the jumps even though red options jumped four times more often.
Methodology
Participants played a board game with six options whose win and loss probabilities changed abruptly without warning, for roughly 500 trials. The authors fitted a Bayesian model that separately tracks risk, estimation uncertainty and the chance of a sudden change, and compared it with model-free reinforcement-learning models. They also tested whether estimation uncertainty acted as a bonus that encourages exploring or a penalty that discourages it, and ran three treatments that varied how much of the task structure participants were told.
Limitations
Key statistics (exact p-values, some group sizes) are missing from the extracted text, so the size of the model differences cannot be checked here. The sample is undergraduates at a single university, and participants in the first treatment could opt into the others, so treatments were not independent random groups. Better model fit shows the Bayesian account describes choices well, not that the brain literally computes these quantities; the neural claims in the discussion come from other studies, not new data.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
People can behave as if they track distinct kinds of uncertainty, and tend to avoid the unknown.
When undergraduates were fully informed that reward probabilities could jump, a Bayesian model that tracks unexpected change fitted their choices better than reinforcement learning, and people penalised options with high estimation uncertainty rather than exploring them.
Evidence for the claim as stated.
Bayesian volatility tracking is not universal: when participants were not told probabilities could jump, the Bayesian model no longer beat reinforcement learning, and in one reversal dataset about 30% of people were best fit by a plain Kalman filter.
Evidence for the claim as stated.
Open questions
Tensions this paper is part of
From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.
Bayesian volatility tracking is not universal: when participants were not told probabilities could jump, the Bayesian model no longer beat reinforcement learning, and in one reversal dataset about 30% of people were best fit by a plain Kalman filter.
Related papers in this topic
Same topic cluster — not a recommendation engine.