Skip to content
PaperFren

Reinforcement learning · Planning · Habits

The standard test of planning versus habit barely rewards planning, and can be passed without it

Evidence: EmergingMore than one study points the same way, but the body is still thin. What the labels mean

Study published Jan 1, 2016. PaperFren added this explanation Sep 26, 2026.

Save this development or follow its topic to track what changes.

Short answer

How much people plan depends on whether planning pays, and the classic two-step task offers almost no payoff while its planning signature can be mimicked.

What happened

One paper found that across nearly all parameters, more model-based control did not raise reward in the original two-step task; redesigning five features made planning pay, and in 381 online participants the planning weight correlated with reward in the new task (r = 0.55) but not the original (r = 0.10). A simulation study found model-free and model-based agents earned rewards on about 0.558 and 0.557 of trials, and that in a reduced task a model-free agent could show the interaction thought to signal planning. A third model treats planning as worth it only when better information beats the reward lost to waiting, reproducing habit formation after overtraining.

Why it matters

Much of the evidence for separate planning and habit systems in humans comes from this task. If planning gives little benefit there, low planning may be rational cost-saving rather than a deficit, and the behavioural signature needs careful checks before it is read as proof of planning.

Evidence

Study type
Agent simulations plus one online behavioural experiment
Sample
381 Mechanical Turk participants; otherwise simulated agents
Journal
PLoS Computational Biology · peer reviewed
Replication
Simulation results are reproducible by construction; the human cost-benefit finding comes from a single experiment
Limitations
The task comparison is correlational and confounded by other design differences; effort cost was not measured directly; the speed/accuracy model is fitted qualitatively to published figures.

What this connects to

Sources

The 3 studies this explanation is built from, by the role each plays. Every source links to PaperFren’s explanation of it and to the original paper.

Primary study

  • Does planning ahead actually earn more reward in lab tasks?

    Kool W, Cushman FA, Gershman SJ · 2016 · PLoS computational biology · 158 citations

    In the most widely used planning-versus-habit task, planning earns essentially no extra reward, but a redesigned task makes planning pay off and people plan more in it.

    What it does not show

    The human result is correlational between two task groups, so the higher planning in the new task could reflect other features such as deterministic transitions or negative rewards rather than a cost-benefit calculation, as the authors note. The paper does not measure effort or cognitive cost directly, only reward. Participants were online workers with limited trials, and model comparison for the original task partly favoured a pure model-free model.

    PaperFren explanationStudy with cards and a quizOriginal paper (DOI)cc by

Supporting evidence

  • Can habits look like planning in the two-step task?

    Akam T, Costa R, Dayan P · 2015 · PLoS computational biology · 116 citations

    Agents that never plan can still produce the behavioural pattern usually taken as proof of planning, so the standard test for 'model-based' choice can be fooled.

    What it does not show

    Everything is simulation; the paper does not show that real humans or animals actually use reward-as-cue or latent-state strategies. The authors argue the artefact is very weak in the original task, so existing human findings are probably not undermined. Model comparison worked on huge simulated datasets, but real datasets are much smaller, real subjects may mix strategies, and fitted models will not match them exactly.

    PaperFren explanationStudy with cards and a quizOriginal paper (DOI)cc by

  • When is it worth stopping to think before acting?

    Keramati M, Dezfouli A, Piray P · 2011 · PLoS computational biology · 255 citations

    A model in which the brain only plans when the expected gain from better information beats the reward lost by waiting explains why well-practised actions become fast, inflexible habits.

    What it does not show

    The fits to data are qualitative comparisons with published figures, not quantitative fits to raw data. The model assumes the planner has perfect knowledge of values, which the authors admit fails in tasks like reversal learning, and it predicts a linear rather than the observed logarithmic growth of reaction time with the number of choices. The proposed role of tonic dopamine as the average reward signal is a hypothesis, not tested here.

    PaperFren explanationStudy with cards and a quizOriginal paper (DOI)cc by

Before

The two-step task cleanly measures a fixed tendency to plan, and its transition-by-outcome pattern is a reliable marker of model-based control.

Now

Planning in that task barely changes reward, people planned more when a redesigned task made it pay, and the marker can be produced by non-planning agents. The human result is a correlation between task groups, most work here is simulation, and the authors argue the artefact is weak in the original task.