Skip to content
PaperFren

Reinforcement learning

Are habits chunked action sequences run by a goal-directed boss?

Dezfouli A, Balleine BW · PLoS computational biology · 2013

Open access · cc by · source: Europe PMC

People's habitual choices were better explained as whole pre-packaged action sequences chosen by a goal-directed system than as separate single actions valued by a model-free habit system.

Study at a glance

Design
Human experiment — Within-subject two-stage decision task with probabilistic transitions and drifting rewards; behaviour analysed with mixed-effects logistic regression and fitted with flat vs hierarchical RL model families.
N
N=15 · 15 adult participants, each completing 270 trials of the two-stage task.
Population
Healthy adult human participants.
Outcome
Probability of repeating first- and second-stage actions, second-stage reaction times, and model-family exceedance probability.

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

First-stage choices mixed habitual and goal-directed patterns, as in earlier work. Crucially, after a rewarded trial, people who repeated their first action also tended to repeat their second action even on a different slot machine, which the hierarchical model predicts and the flat model does not. These sequence-completing responses were faster than second actions made outside a sequence, and the hierarchical model family beat the flat family with an exceedance probability of about 0.99.

Methodology

Fifteen participants played a two-stage task: a first key press led (usually, but not always) to one of two slot machines, where a second key press could win money with slowly changing probabilities. The authors contrasted a 'flat' account, where model-free habits and model-based planning compete at every step, with a 'hierarchical' account, where a goal-directed system can choose either single actions or two-step action sequences that then run without checking feedback. They tested the accounts' diverging predictions about second-stage choices and reaction times and compared model families by Bayesian model selection.

Limitations

The sample was only 15 people on a single modified task, so the finding needs replication at scale and on the original symbol-based version of the task, where second-stage actions cannot be directly compared. The results cannot tell whether action sequences are themselves valued in a goal-directed or model-free way, and a flat system running in parallel cannot be ruled out, only shown to be unnecessary. Both model families also missed two features of the data that the authors chose not to model.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • Habit vs plan may be a spectrum built from shared parts.

    Intermediate or hierarchical mechanisms can explain behaviour sitting between the two systems: people repeated whole rewarded action sequences (hierarchical models beat flat ones, exceedance probability ≈0.99), and successor-representation agents with enough offline replay solved revaluation tasks like a full planner but fell back to cached behaviour with too little replay.

    Evidence for the claim as stated.

  • Whether the planning signature in existing human two-step data is trustworthy: one simulation shows model-free artefacts can mimic it, but its authors argue the artefact is very weak in the original task, while the action-sequence study shows hierarchical habits can masquerade as planning in human data.

    Evidence for the claim as stated.

Open questions

Tensions this paper is part of

From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.

Related papers in this topic

Same topic cluster — not a recommendation engine.