Skip to content
PaperFren

Uncertainty and calibration

How should a learner speed up learning when the world keeps changing?

Piray P, Daw ND · PLoS computational biology · 2020

Open access · cc by · source: Europe PMC

A simple extension of the Kalman filter that also tracks how fast the world is changing approximates optimal learning more accurately than the popular HGF and explains most people's choices better.

Study at a glance

Design
Computational / modelling — New Bayesian learning algorithm (VKF) tested in simulations against ground truth, a particle filter and the HGF, then fitted to two existing human choice datasets with hierarchical model comparison.
N
Two reanalysed human datasets: 44 participants in the first task and 161 analysed (of 174) in the second; simulations used 1000 or 500 generated sequences.
Population
Simulated volatile environments; adult participants in two published probabilistic reversal-learning tasks.
Outcome
Tracking error relative to a near-optimal particle filter; learning-rate/volatility relationship; model evidence for fitting human choices.

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

Measured as extra error over the particle-filter benchmark, the VKF was off by only a few percent while the HGF was off by about a fifth, and the HGF sometimes crashed with impossible negative variances. In the binary version, VKF's learning rate rose with its volatility estimate (as Pearce-Hall theory predicts), whereas the binary HGF showed the opposite relationship in almost all simulations. VKF was the best-fitting model for 37 of 44 people in the first dataset and 102 of 161 in the second, and most people's fitted learning rates rose after contingency switches.

Methodology

The authors built a generative model in which both a hidden reward value and its rate of change (volatility) drift over time, and derived a cheap approximate learning rule for it called the volatile Kalman filter (VKF), plus a version for binary win/lose outcomes. They compared its estimates with a near-exact particle filter and with the hierarchical Gaussian filter (HGF) on simulated data. They then fitted VKF, HGF, a plain Kalman filter and a Rescorla-Wagner model to choices from two existing human reversal-learning datasets.

Limitations

The human tests reuse two existing datasets rather than a task designed to discriminate the models, and in the second task about 30% of people were best fit by a plain Kalman filter, so volatility tracking is not universal. VKF only handles two levels and gives up the HGF's ability to stack deeper hierarchies. The observation noise is treated as a fixed parameter, whereas real learners may need to estimate it too, and parameter recovery showed the volatility parameters were harder to estimate reliably. Better model fit does not prove the brain literally runs this algorithm.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • Learning rates that rise with estimated volatility describe many people's behaviour.

    A two-level volatile Kalman filter came within a few percent of a particle-filter benchmark while the hierarchical Gaussian filter erred by about a fifth, and the volatile Kalman filter was the best-fitting model for 37 of 44 and 102 of 161 participants in two reversal-learning datasets.

    Evidence for the claim as stated.

  • Bayesian volatility tracking is not universal: when participants were not told probabilities could jump, the Bayesian model no longer beat reinforcement learning, and in one reversal dataset about 30% of people were best fit by a plain Kalman filter.

    Evidence for the claim as stated.

Open questions

Tensions this paper is part of

From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.

Related papers in this topic

Same topic cluster — not a recommendation engine.