Reinforcement learning · Planning · Habits
The standard test of planning versus habit barely rewards planning, and can be passed without it
Save this development or follow its topic to track what changes.
Short answer
How much people plan depends on whether planning pays, and the classic two-step task offers almost no payoff while its planning signature can be mimicked.
What happened
One paper found that across nearly all parameters, more model-based control did not raise reward in the original two-step task; redesigning five features made planning pay, and in 381 online participants the planning weight correlated with reward in the new task (r = 0.55) but not the original (r = 0.10). A simulation study found model-free and model-based agents earned rewards on about 0.558 and 0.557 of trials, and that in a reduced task a model-free agent could show the interaction thought to signal planning. A third model treats planning as worth it only when better information beats the reward lost to waiting, reproducing habit formation after overtraining.
Why it matters
Much of the evidence for separate planning and habit systems in humans comes from this task. If planning gives little benefit there, low planning may be rational cost-saving rather than a deficit, and the behavioural signature needs careful checks before it is read as proof of planning.
Evidence
- Study type
- Agent simulations plus one online behavioural experiment
- Sample
- 381 Mechanical Turk participants; otherwise simulated agents
- Journal
- PLoS Computational Biology · peer reviewed
- Replication
- Simulation results are reproducible by construction; the human cost-benefit finding comes from a single experiment
- Limitations
- The task comparison is correlational and confounded by other design differences; effort cost was not measured directly; the speed/accuracy model is fitted qualitatively to published figures.
What this connects to
Sources
The 3 studies this explanation is built from, by the role each plays. Every source links to PaperFren’s explanation of it and to the original paper.
Primary study
- Does planning ahead actually earn more reward in lab tasks?
In the most widely used planning-versus-habit task, planning earns essentially no extra reward, but a redesigned task makes planning pay off and people plan more in it.
What it does not showLimitations
The human result is correlational between two task groups, so the higher planning in the new task could reflect other features such as deterministic transitions or negative rewards rather than a cost-benefit calculation, as the authors note. The paper does not measure effort or cognitive cost directly, only reward. Participants were online workers with limited trials, and model comparison for the original task partly favoured a pure model-free model.
PaperFren explanationStudy with cards and a quizOriginal paper (DOI)cc by
Supporting evidence
- Can habits look like planning in the two-step task?
Agents that never plan can still produce the behavioural pattern usually taken as proof of planning, so the standard test for 'model-based' choice can be fooled.
What it does not showLimitations
Everything is simulation; the paper does not show that real humans or animals actually use reward-as-cue or latent-state strategies. The authors argue the artefact is very weak in the original task, so existing human findings are probably not undermined. Model comparison worked on huge simulated datasets, but real datasets are much smaller, real subjects may mix strategies, and fitted models will not match them exactly.
PaperFren explanationStudy with cards and a quizOriginal paper (DOI)cc by
- When is it worth stopping to think before acting?
A model in which the brain only plans when the expected gain from better information beats the reward lost by waiting explains why well-practised actions become fast, inflexible habits.
What it does not showLimitations
The fits to data are qualitative comparisons with published figures, not quantitative fits to raw data. The model assumes the planner has perfect knowledge of values, which the authors admit fails in tasks like reversal learning, and it predicts a linear rather than the observed logarithmic growth of reaction time with the number of choices. The proposed role of tonic dopamine as the average reward signal is a hypothesis, not tested here.
PaperFren explanationStudy with cards and a quizOriginal paper (DOI)cc by
Before
The two-step task cleanly measures a fixed tendency to plan, and its transition-by-outcome pattern is a reliable marker of model-based control.
Now
Planning in that task barely changes reward, people planned more when a redesigned task made it pay, and the marker can be produced by non-planning agents. The human result is a correlation between task groups, most work here is simulation, and the authors argue the artefact is weak in the original task.