Skip to content
PaperFren

Do people learn like Bayesians when rewards suddenly change?

Open paper intelligence

People chose like Bayesian learners who dislike uncertainty, but only when told that reward odds could suddenly jump; otherwise simple trial-and-error learning fit just as well.

Source

Risk, unexpected uncertainty, and estimation uncertainty: Bayesian learning in unstable settings

Payzan-LeNestour E, Bossaerts P · PLoS computational biology · 2011

doi.org/10.1371/journal.pcbi.1001048Read the full paper ↗173 citationscc by

Study at a glance

Design
Human experiment — Model comparison (Bayesian learner with forgetting vs Rescorla-Wagner and Pearce-Hall) on choices from a six-arm restless bandit, plus a new experiment manipulating how much task structure participants were told.
N
43 undergraduates in the no-information treatment; the numbers in the other two treatments are garbled in the extracted text, and treatments 1 and 2 pooled gave 75 cases. Reanalysed data from an earlier experiment are also used.
Population
Undergraduates at Ecole Polytechnique Fédérale Lausanne
Outcome
Model fit (log-likelihood, BIC) of learning models to trial-by-trial choices; debriefing reports of noticing jumps

Structured fields used in claim comparison tables when every cited study has a complete layer.

What they did

Participants played a board game with six options whose win and loss probabilities changed abruptly without warning, for roughly 500 trials. The authors fitted a Bayesian model that separately tracks risk, estimation uncertainty and the chance of a sudden change, and compared it with model-free reinforcement-learning models. They also tested whether estimation uncertainty acted as a bonus that encourages exploring or a penalty that discourages it, and ran three treatments that varied how much of the task structure participants were told.

What they found

When fully informed about the task, participants' choices were better fitted by the Bayesian model than by reinforcement learning, and the Pearce-Hall model fitted worst. Adding a penalty for estimation uncertainty improved the fit, while an exploration bonus made it worse, so people avoided rather than sought uncertain options. When participants were not told that probabilities could jump, the Bayesian model no longer beat reinforcement learning, and debriefing showed most had not noticed the jumps even though red options jumped four times more often.

The limits

What it doesn't show

Key statistics (exact p-values, some group sizes) are missing from the extracted text, so the size of the model differences cannot be checked here. The sample is undergraduates at a single university, and participants in the first treatment could opt into the others, so treatments were not independent random groups. Better model fit shows the Bayesian account describes choices well, not that the brain literally computes these quantities; the neural claims in the discussion come from other studies, not new data.

Key terms

Restless bandit
A choice task with several options whose reward probabilities change over time, so the best option must be relearned.
Risk
Uncertainty that remains even when the outcome probabilities are known exactly.
Estimation uncertainty (ambiguity)
Uncertainty about what the outcome probabilities are, which shrinks as you gather data.
Unexpected uncertainty
The chance that the environment has suddenly changed, meaning old learning should be discarded.
Model-free reinforcement learning
Learning option values directly from prediction errors without a model of how outcomes are generated.
Softmax
A rule that turns option values into choice probabilities, balancing exploiting the best option and exploring others.

Flashcards

1 / 10

0 of 10 answers reviewed

Research intelligence for this paper

See its role on concept claims, tensions it is part of, placement history, and related discoveries.

Open paper intelligence

Quiz yourself

1 / 5

Which kind of uncertainty signals that past learning should be partly forgotten?

Common questions

Why does it matter to separate three kinds of uncertainty?

Each should change learning differently: risk should be ignored, estimation uncertainty calls for more learning, and a likely sudden change means you should start learning again.

Did people explore uncertain options more?

No. The model fitted better when uncertainty lowered an option's value, consistent with ambiguity aversion rather than an exploration bonus.

Why did uninformed participants look like simple reinforcement learners?

Most did not notice the sudden jumps and blamed bad runs on chance, so they had no model of change to use; the authors suggest people fall back on model-free learning under structural uncertainty.

More on Uncertainty and calibration