Skip to content
PaperFren

Reinforcement learning

Do teenagers learn from rewards and punishments like adults?

Palminteri S, Kilford EJ, Coricelli G, et al. · PLoS computational biology · 2016

Open access · cc by · source: Europe PMC

Adolescents' learning was best explained by a simple reward-tracking algorithm, while adults also learned from outcomes they didn't choose and from context, which helped them avoid punishments.

Study at a glance

Design
Human experiment — Within-subject 2x2 probabilistic learning task (reward vs punishment, partial vs complete feedback) with computational model comparison between age groups
N
N=38 · 38 participants after IQ matching (18 adolescents aged 12-17, 20 adults aged 18-32), from 50 recruited
Population
Adolescents and adults recruited in London
Outcome
Model posterior probabilities, correct-choice rate, reaction times and post-learning cue preferences

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

Model comparison favoured basic Q-learning in adolescents and the full three-module model in adults, with a strong group-by-model interaction. Behaviourally, adolescents improved less at avoiding punishment and did not benefit from seeing the unchosen outcome, while adults did. Both groups learned equally well in the simplest reward-with-partial-feedback context, arguing against a general motivation or attention deficit.

Methodology

Adolescents and adults chose between pairs of symbols that led to winning or losing points with 75% or 25% probability. Some pairs offered rewards and others threatened losses, and some pairs showed only the chosen outcome while others also showed the outcome of the option not chosen. The authors fitted three nested reinforcement-learning models (basic Q-learning, plus counterfactual learning, plus value contextualisation) to each person's choices and compared them, then checked predictions against behaviour and a later test.

Limitations

The final sample was small, 38 people, after excluding participants to match non-verbal IQ, and the design is cross-sectional, so it cannot track change within individuals. The age cut-off at 18 is arbitrary and the adolescent range is wide, which the authors acknowledge. Some behavioural effects (such as the feedback-by-group interaction on choice rate) were not statistically significant, and links to brain development are inferred from other studies rather than measured.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • Adolescents use a simpler learning strategy, not less motivation.

    Learning strategies change with development: basic Q-learning fitted adolescents best while adults were best fit by a model including punishment-context and counterfactual modules; adolescents did not benefit from seeing unchosen outcomes, though both groups learned equally in the simplest context.

    Evidence for the claim as stated.

  • Whether punishment carries a real learning signal: in the approach/withdrawal task punishment's average learning effect was indistinguishable from zero, while adults in the adolescence study used punishment-context information that adolescents lacked.

    Same question, contrary or null result.

Open questions

Tensions this paper is part of

From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.

Related papers in this topic

Same topic cluster — not a recommendation engine.