Concept · artificial-intelligence
Model-based vs model-free control (goal-directed vs habitual)
5 studies1 discoveryEvidence last moved Sep 27, 2026
Model-based control plans by using a learned map of how actions lead to states and rewards; model-free control just caches which actions paid off before (habits). The evidence here combines computational simulations of the widely used two-step task and of hybrid algorithms with small human experiments.
The two-step task is used to link planning to psychiatric traits and brain regions, so students need to know what its signatures do and don't prove. These papers show that planning often doesn't pay in the standard task, that habits can mimic planning, and that the two systems may not be cleanly separate.
Studies
5
Findings
5
7 supporting · 0 challenging · 1 qualifying citations
Open tensions
1
Latest change
Concept page published
Model-based vs model-free control (goal-directed vs habitual)
Currently
What we know
- The classic task gives little incentive to plan.
- People may use planning more when it is worth the effort.
- A 'planning signature' is not proof of planning.
- Habit vs plan may be a spectrum built from shared parts.
- Habits can be seen as the fast option once planning stops being worth its time.
Largest unresolved question
Whether the planning signature in existing human two-step data is trustworthy: one simulation shows model-free artefacts can mimic it, but its authors argue the artefact is very weak in the original task, while the action-sequence study shows hierarchical habits can masquerade as planning in human data.
Common misconceptions
A high model-based weight on the two-step task means a person earns more by being smarter.
In the original task planning barely changes reward; only redesigned tasks link model-based weight to reward.
Model-based and model-free are two sealed-off systems in the brain.
Successor-representation and action-sequence models show hybrids that fit or explain the data; brain mappings in these papers are hypotheses, not measurements.
Related
Claim ledger
What the evidence shows
Drawn from 5 studies in this library. Mix labels say which citation roles are present; they are not a strength score. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope.
The classic task gives little incentive to plan.
In the original two-step task, planning barely earns more reward: simulated model-free and model-based agents were rewarded on about 0.558 and 0.557 of trials, and more model-based control did not raise reward across nearly the whole parameter range of the original task and two variants.
- Can habits look like planning in the two-step task?
- Does planning ahead actually earn more reward in lab tasks?
Study Role Design N Population Outcome Can habits look like planning in the two-step task? Supports Computational / modellingSimulated model-free, model-based, reward-as-cue and latent-state agents on the original and a reduced two-step task, analysed with stay-probability regressions, lagged regressions and likelihood model comparison. No participants; results come from large simulated datasets of agent choices. Simulated reinforcement-learning agents. Fraction of rewarded trials; regression loading on the transition-by-outcome interaction; model-fit likelihood. Does planning ahead actually earn more reward in lab tasks? Supports Human experimentMonte Carlo simulations of hybrid model-based/model-free RL agents across task variants, then an online experiment assigning participants to the original Daw two-step task or a new task, with RL model fitting. N=381 · 381 Mechanical Turk participants analysed after exclusions, split between the novel task and the original Daw task; simulations used synthetic agents. Adult online participants recruited through Amazon Mechanical Turk; simulated reinforcement-learning agents. Relationship between the model-based weighting parameter w and reward rate; average w in each task. People may use planning more when it is worth the effort.
When a redesigned task made planning pay, people planned more: model-based weight correlated with reward in the new task (r = 0.55) but not the original (r = 0.10), and median weight was higher (0.48 vs 0.27).
- Does planning ahead actually earn more reward in lab tasks?— Between-group correlational comparison of online workers; other task features could explain the difference, and effort was not measured.
A 'planning signature' is not proof of planning.
The behavioural signature of planning can be produced without planning: a model-free agent showed the transition-by-outcome interaction on a reduced task, and a latent-state agent looked almost identical to a planner on regression tests, only separable by likelihood model comparison.
Habit vs plan may be a spectrum built from shared parts.
Intermediate or hierarchical mechanisms can explain behaviour sitting between the two systems: people repeated whole rewarded action sequences (hierarchical models beat flat ones, exceedance probability ≈0.99), and successor-representation agents with enough offline replay solved revaluation tasks like a full planner but fell back to cached behaviour with too little replay.
- Are habits chunked action sequences run by a goal-directed boss?
- Can a predictive map explain planning with simple learning?
Study Role Design N Population Outcome Are habits chunked action sequences run by a goal-directed boss? Supports Human experimentWithin-subject two-stage decision task with probabilistic transitions and drifting rewards; behaviour analysed with mixed-effects logistic regression and fitted with flat vs hierarchical RL model families. N=15 · 15 adult participants, each completing 270 trials of the two-stage task. Healthy adult human participants. Probability of repeating first- and second-stage actions, second-stage reaction times, and model-family exceedance probability. Can a predictive map explain planning with simple learning? Supports Computational / modellingGrid-world simulations of SR-TD, SR-MB and SR-Dyna agents (plus model-free and Dyna-Q foils) on latent learning, detour and policy revaluation tasks. No sample size: results are simulated agent choices on the first test trial after each task manipulation; the Methods section is not in the available text. Simulated reinforcement-learning agents in grid-world mazes modelled on Tolman's rodent experiments Whether each agent immediately chooses the correct new shortest path after a reward change, a transition change, or a policy-relevant reward change Habits can be seen as the fast option once planning stops being worth its time.
A speed/accuracy account reproduces the classic overtraining effect: a simulated agent stopped deliberating after about 100 training trials and kept responding for a devalued outcome after extensive training, but stayed goal-directed when two equally valued outcomes were available.
Debates
Tensions and limits
Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes.
Whether the planning signature in existing human two-step data is trustworthy: one simulation shows model-free artefacts can mimic it, but its authors argue the artefact is very weak in the original task, while the action-sequence study shows hierarchical habits can masquerade as planning in human data.
Whether the planning signature in existing human two-step data is trustworthy: one simulation shows model-free artefacts can mimic it, but its authors argue the artefact is very weak in the original task, while the action-sequence study shows hierarchical habits can masquerade as planning in human data.
- Can habits look like planning in the two-step task?
- Are habits chunked action sequences run by a goal-directed boss?
Study Role Design N Population Outcome Can habits look like planning in the two-step task? Supports Computational / modellingSimulated model-free, model-based, reward-as-cue and latent-state agents on the original and a reduced two-step task, analysed with stay-probability regressions, lagged regressions and likelihood model comparison. No participants; results come from large simulated datasets of agent choices. Simulated reinforcement-learning agents. Fraction of rewarded trials; regression loading on the transition-by-outcome interaction; model-fit likelihood. Are habits chunked action sequences run by a goal-directed boss? Supports Human experimentWithin-subject two-stage decision task with probabilistic transitions and drifting rewards; behaviour analysed with mixed-effects logistic regression and fitted with flat vs hierarchical RL model families. N=15 · 15 adult participants, each completing 270 trials of the two-stage task. Healthy adult human participants. Probability of repeating first- and second-stage actions, second-stage reaction times, and model-family exceedance probability.
PaperFren reads this as a limit on how far one study travels — different assays, populations, or outcomes — not a forced fight between papers.
Timeline
How understanding moved
Study years are when the paper was published. Evidence edits are dated changes to this page's claims. Explanations are when PaperFren added a Discovery — not a claim that the science happened that day.
2026
- The standard test of planning versus habit barely rewards planning, and can be passed without it
Concept page published
Model-based vs model-free control (goal-directed vs habitual)
Change log
What changed
Dated edits to this page's evidence: studies added or removed from a claim, claims added or withdrawn, and new explanations tagged here. Rewordings are not listed.
- Concept page published
Papers
5 studies in this library bear on Model-based vs model-free control (goal-directed vs habitual), ordered by citations.
- When is it worth stopping to think before acting?
A model in which the brain only plans when the expected gain from better information beats the reward lost by waiting explains why well-practised actions become fast, inflexible habits.
- Can a predictive map explain planning with simple learning?
Learning values on top of a map of which states tend to follow which lets an agent show flexible, planning-like behaviour using the same prediction-error rule linked to dopamine, and adding offline replay closes the remaining gaps.
- Does planning ahead actually earn more reward in lab tasks?
In the most widely used planning-versus-habit task, planning earns essentially no extra reward, but a redesigned task makes planning pay off and people plan more in it.
- Are habits chunked action sequences run by a goal-directed boss?
People's habitual choices were better explained as whole pre-packaged action sequences chosen by a goal-directed system than as separate single actions valued by a model-free habit system.
- Can habits look like planning in the two-step task?
Agents that never plan can still produce the behavioural pattern usually taken as proof of planning, so the standard test for 'model-based' choice can be fooled.
Compare studies
Select 2–10 studies. Design and N are labels, not a ranking.
Nothing selected yet.
Questions
What is still open
Whether the planning signature in existing human two-step data is trustworthy: one simulation shows model-free artefacts can mimic it, but its authors argue the artefact is very weak in the original task, while the action-sequence study shows hierarchical habits can masquerade as planning in human data.
Ask PaperFren about Model-based vs model-free control (goal-directed vs habitual)
Study this conceptflashcards and short-answer questions
Why might the original two-step task underestimate people's capacity to plan?
Simulations show model-based and model-free agents earn almost identical reward in the original task, so there is little incentive to plan. Features such as slowly drifting, low-contrast reward probabilities remove the benefit. When a redesigned task made planning pay, participants showed higher model-based weights, correlated with reward, though that comparison is correlational.
What does 'SR-Dyna' show about the relation between habits and planning?
In grid-world simulations, a successor-representation agent updated by offline replay solved latent learning, detour and policy revaluation tasks like full planning, but with little replay it fell back to cached SR-TD behaviour. This suggests planning-like flexibility can come from cached predictive maps plus replay. It is a simulation, not a fit to animal data.