How do people learn whether an adviser is trying to help them?
People's choices were best explained by a learning model that tracks not only how accurate an adviser is but also how quickly the adviser's intentions are changing.
Source
Inferring on the intentions of others by hierarchical Bayesian learning
Study at a glance
- Design
- Human experiment — Pairs of men played an incentivised advice game (social task) plus a blindfolded-adviser control task; 12 learning/response models fitted per player and compared with random-effects Bayesian model selection.
- N
- N=16 · 32 male volunteers formed 16 player-adviser pairs; models were fitted to the 16 players' choices across 192 trials per task.
- Population
- Healthy adult men aged 19-30 in Zurich
- Outcome
- Which computational model best explains players' trial-wise decisions to follow or reject advice; links between model parameters, ratings, questionnaires and performance
Structured fields used in claim comparison tables when every cited study has a complete layer.
What they did
Pairs of male volunteers played a lottery game in which a player predicted a colour from a pie chart while an adviser, who knew more, held up a card recommending a colour. The adviser's pay-off secretly rewarded helping at some points and misleading at others, so his intentions shifted over the game. The researchers fitted 12 candidate models (three learning models crossed with four ways of turning beliefs into choices) to each player's choices and compared them with Bayesian model selection. A control task used a blindfolded adviser drawing cards from decks, removing any intention.
What they found
A three-level hierarchical Gaussian filter, in which the learning rate rises when the adviser seems volatile, beat simpler models including Rescorla-Wagner, and choices became noisier when players believed the adviser's intentions were unstable. Players combined the advice with the pie chart but weighted the pie chart more. The winning model's parameters predicted perspective-taking questionnaire scores and task accuracy, and its trial-by-trial estimates tracked players' explicit ratings of the adviser, whereas the Rescorla-Wagner estimates did not. In the non-intentional control task players did worse and relied less on the advice.
The limits
What it doesn't show
The sample is small (16 players) and all male, chosen deliberately to avoid gender effects, so it cannot say whether women or larger, more varied groups learn the same way. Model selection only shows which of the 12 tested models fits best, not that the brain literally implements a hierarchical Gaussian filter; untested models could fit better. The authors note they modelled only the players and not the advisers, and they did not model recursive reasoning (thinking about what the other thinks about you). Advisers chose their own strategies, so the volatility each player faced was not experimentally controlled.
Key terms
- Hierarchical Gaussian filter (HGF)
- A Bayesian learning model with stacked levels, where a higher level estimates how volatile the environment is and so controls how fast the lower level updates its beliefs.
- Rescorla-Wagner model
- A classic reinforcement-learning rule that updates a value estimate by a fixed learning rate times the prediction error.
- Volatility
- How quickly the hidden state of the world, here the adviser's intentions, changes over time.
- Bayesian model selection
- Comparing candidate models by their approximate model evidence, which balances fit to the data against model complexity.
- Exceedance probability
- The probability that one model (or family of models) is more common in the population than every other model compared.
- Precision-weighted prediction error
- A surprise signal scaled by how reliable the new input is relative to the current belief; it acts like a dynamic learning rate.
Flashcards
0 of 10 answers reviewed
Research intelligence for this paper
See its role on concept claims, tensions it is part of, placement history, and related discoveries.
Quiz yourself
What makes the HGF different from the Rescorla-Wagner model?
Common questions
Why not just use a normal reinforcement-learning model?
Rescorla-Wagner uses one fixed learning rate, which is too fast when the adviser is stable and too slow when he suddenly switches. The hierarchical model adjusts its learning rate trial by trial, and it fitted players' choices and ratings better.
What was the control task for?
It kept the same volatility and advice structure but removed intention, because a blindfolded adviser drew cards from decks. That let the authors check which effects were specific to reading another person's intentions.
Does the model predict anything beyond the choices it was fitted to?
Yes. Its parameters predicted empathy-related questionnaire scores collected days earlier and overall accuracy, and its belief estimates matched players' explicit ratings of the adviser.
More on Reinforcement learning