Skip to content
PaperFren

Reinforcement learning

Can habits look like planning in the two-step task?

Akam T, Costa R, Dayan P · PLoS computational biology · 2015

Open access · cc by · source: Europe PMC

Agents that never plan can still produce the behavioural pattern usually taken as proof of planning, so the standard test for 'model-based' choice can be fooled.

Study at a glance

Design
Computational / modelling — Simulated model-free, model-based, reward-as-cue and latent-state agents on the original and a reduced two-step task, analysed with stay-probability regressions, lagged regressions and likelihood model comparison.
N
No participants; results come from large simulated datasets of agent choices.
Population
Simulated reinforcement-learning agents.
Outcome
Fraction of rewarded trials; regression loading on the transition-by-outcome interaction; model-fit likelihood.

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

On the original task, planning barely paid: the model-free and model-based agents earned rewards on about 0.558 and 0.557 of trials. On the reduced task planning paid more (0.649 versus 0.594), but the model-free agent now showed the transition-by-outcome interaction thought to signal planning, because its starting action values correlated with later trial events; adding a 'repeat correct choice' predictor removed this artefact. The latent-state agent looked almost identical to a planner on regression tests and could only be told apart by likelihood model comparison.

Methodology

The authors simulated reinforcement-learning agents on the popular two-step task and on a reduced version designed for animals, with more common transitions (probability raised from 0.7 to 0.8), no second-step choice and blocks of clearly good and bad options. They compared a simple model-free agent, a planning (model-based) agent, and two agents that treat where the last reward came from as a cue, including one that infers a hidden 'which side is good' state. They analysed choices with the usual stay-probability test, an extended regression and model fitting.

Limitations

Everything is simulation; the paper does not show that real humans or animals actually use reward-as-cue or latent-state strategies. The authors argue the artefact is very weak in the original task, so existing human findings are probably not undermined. Model comparison worked on huge simulated datasets, but real datasets are much smaller, real subjects may mix strategies, and fitted models will not match them exactly.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • The classic task gives little incentive to plan.

    In the original two-step task, planning barely earns more reward: simulated model-free and model-based agents were rewarded on about 0.558 and 0.557 of trials, and more model-based control did not raise reward across nearly the whole parameter range of the original task and two variants.

    Evidence for the claim as stated.

  • A 'planning signature' is not proof of planning.

    The behavioural signature of planning can be produced without planning: a model-free agent showed the transition-by-outcome interaction on a reduced task, and a latent-state agent looked almost identical to a planner on regression tests, only separable by likelihood model comparison.

    Evidence for the claim as stated.

  • Whether the planning signature in existing human two-step data is trustworthy: one simulation shows model-free artefacts can mimic it, but its authors argue the artefact is very weak in the original task, while the action-sequence study shows hierarchical habits can masquerade as planning in human data.

    Evidence for the claim as stated.

Open questions

Tensions this paper is part of

From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.

Discoveries this paper informs or conflicts with

Related papers in this topic

Same topic cluster — not a recommendation engine.