Reinforcement learning
How do people learn whether an adviser is trying to help them?
Open access · cc by · source: Europe PMC
People's choices were best explained by a learning model that tracks not only how accurate an adviser is but also how quickly the adviser's intentions are changing.
Study at a glance
- Design
- Human experiment — Pairs of men played an incentivised advice game (social task) plus a blindfolded-adviser control task; 12 learning/response models fitted per player and compared with random-effects Bayesian model selection.
- N
- N=16 · 32 male volunteers formed 16 player-adviser pairs; models were fitted to the 16 players' choices across 192 trials per task.
- Population
- Healthy adult men aged 19-30 in Zurich
- Outcome
- Which computational model best explains players' trial-wise decisions to follow or reject advice; links between model parameters, ratings, questionnaires and performance
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
A three-level hierarchical Gaussian filter, in which the learning rate rises when the adviser seems volatile, beat simpler models including Rescorla-Wagner, and choices became noisier when players believed the adviser's intentions were unstable. Players combined the advice with the pie chart but weighted the pie chart more. The winning model's parameters predicted perspective-taking questionnaire scores and task accuracy, and its trial-by-trial estimates tracked players' explicit ratings of the adviser, whereas the Rescorla-Wagner estimates did not. In the non-intentional control task players did worse and relied less on the advice.
Methodology
Pairs of male volunteers played a lottery game in which a player predicted a colour from a pie chart while an adviser, who knew more, held up a card recommending a colour. The adviser's pay-off secretly rewarded helping at some points and misleading at others, so his intentions shifted over the game. The researchers fitted 12 candidate models (three learning models crossed with four ways of turning beliefs into choices) to each player's choices and compared them with Bayesian model selection. A control task used a blindfolded adviser drawing cards from decks, removing any intention.
Limitations
The sample is small (16 players) and all male, chosen deliberately to avoid gender effects, so it cannot say whether women or larger, more varied groups learn the same way. Model selection only shows which of the 12 tested models fits best, not that the brain literally implements a hierarchical Gaussian filter; untested models could fit better. The authors note they modelled only the players and not the advisers, and they did not model recursive reasoning (thinking about what the other thinks about you). Advisers chose their own strategies, so the volatility each player faced was not experimentally controlled.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Learning rates are not fixed; they rise when the world seems unstable.
When learning about others, people adjust learning speed to perceived volatility: a hierarchical Gaussian filter beat Rescorla-Wagner in a social advice task, and its trial-by-trial estimates tracked players' explicit ratings of the adviser.
Evidence for the claim as stated.
Related papers in this topic
Same topic cluster — not a recommendation engine.