Skip to content
PaperFren

Are habits chunked action sequences run by a goal-directed boss?

Open paper intelligence

People's habitual choices were better explained as whole pre-packaged action sequences chosen by a goal-directed system than as separate single actions valued by a model-free habit system.

Source

Actions, action sequences and habits: evidence that goal-directed and habitual action control are hierarchically organized

Dezfouli A, Balleine BW · PLoS computational biology · 2013

doi.org/10.1371/journal.pcbi.1003364Read the full paper ↗153 citationscc by

Study at a glance

Design
Human experiment — Within-subject two-stage decision task with probabilistic transitions and drifting rewards; behaviour analysed with mixed-effects logistic regression and fitted with flat vs hierarchical RL model families.
N
N=15 · 15 adult participants, each completing 270 trials of the two-stage task.
Population
Healthy adult human participants.
Outcome
Probability of repeating first- and second-stage actions, second-stage reaction times, and model-family exceedance probability.

Structured fields used in claim comparison tables when every cited study has a complete layer.

What they did

Fifteen participants played a two-stage task: a first key press led (usually, but not always) to one of two slot machines, where a second key press could win money with slowly changing probabilities. The authors contrasted a 'flat' account, where model-free habits and model-based planning compete at every step, with a 'hierarchical' account, where a goal-directed system can choose either single actions or two-step action sequences that then run without checking feedback. They tested the accounts' diverging predictions about second-stage choices and reaction times and compared model families by Bayesian model selection.

What they found

First-stage choices mixed habitual and goal-directed patterns, as in earlier work. Crucially, after a rewarded trial, people who repeated their first action also tended to repeat their second action even on a different slot machine, which the hierarchical model predicts and the flat model does not. These sequence-completing responses were faster than second actions made outside a sequence, and the hierarchical model family beat the flat family with an exceedance probability of about 0.99.

The limits

What it doesn't show

The sample was only 15 people on a single modified task, so the finding needs replication at scale and on the original symbol-based version of the task, where second-stage actions cannot be directly compared. The results cannot tell whether action sequences are themselves valued in a goal-directed or model-free way, and a flat system running in parallel cannot be ruled out, only shown to be unnecessary. Both model families also missed two features of the data that the authors chose not to model.

Key terms

Goal-directed action
Behaviour chosen by predicting its consequences and their current value, using a model of the environment.
Habit
Behaviour repeated because it was rewarded before, insensitive to changes in outcome value or action-outcome contingency.
Model-free reinforcement learning
Learning action values from cached reward history without representing how actions lead to states.
Action sequence (chunk)
Several actions bundled into one unit that, once started, runs to completion without checking intermediate feedback.
Two-stage task
A choice task where a first choice leads probabilistically to a second choice point, used to separate model-based from habitual control.
Exceedance probability
In Bayesian model comparison, the probability that one model (or family) is more common in the population than the others.

Flashcards

1 / 12

0 of 12 answers reviewed

Research intelligence for this paper

See its role on concept claims, tensions it is part of, placement history, and related discoveries.

Open paper intelligence

Quiz yourself

1 / 5

In the hierarchical account, what makes an action habitual?

Common questions

How can habits be explained without a model-free system?

If a rewarded two-step sequence is simply repeated as a unit, the first action looks habitual because the sequence ignores which slot machine it led to.

What was the decisive test between the accounts?

Whether people repeated the second-stage action on a different slot machine after repeating the first action; only the sequence account predicts this.

Why do reaction times matter here?

Open-loop sequences should run faster; slower second responses suggest the sequence was interrupted and control returned to deliberation.

More on Reinforcement learning