Skip to content
PaperFren

Reinforcement learning

Does planning ahead actually earn more reward in lab tasks?

Kool W, Cushman FA, Gershman SJ · PLoS computational biology · 2016

Open access · cc by · source: Europe PMC

In the most widely used planning-versus-habit task, planning earns essentially no extra reward, but a redesigned task makes planning pay off and people plan more in it.

Study at a glance

Design
Human experiment — Monte Carlo simulations of hybrid model-based/model-free RL agents across task variants, then an online experiment assigning participants to the original Daw two-step task or a new task, with RL model fitting.
N
N=381 · 381 Mechanical Turk participants analysed after exclusions, split between the novel task and the original Daw task; simulations used synthetic agents.
Population
Adult online participants recruited through Amazon Mechanical Turk; simulated reinforcement-learning agents.
Outcome
Relationship between the model-based weighting parameter w and reward rate; average w in each task.

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

Across nearly the whole parameter range, more model-based control did not raise reward in the original task or two published variants. Five features restricted the benefit: low-contrast reward probabilities, slow drift, rare transitions, a choice at the second stage, and noisy binary rewards; fixing all five together boosted the planning-reward link by a factor of about 230. In people, w correlated with reward in the new task (r = 0.55) but not the original (r = 0.10), and median w was higher in the new task (0.48 vs 0.27).

Methodology

The authors simulated agents that mix model-free learning (repeat what was rewarded) and model-based planning (use a map of the task), varying the mixing weight w, learning rate and choice randomness, on the Daw two-step task and several variants. They then changed task features one at a time to find what makes planning rewarding, and combined all of them into a new task. Finally, online participants played either the original task or the new task for 125 trials each, and hybrid RL models were fitted to each person's choices.

Limitations

The human result is correlational between two task groups, so the higher planning in the new task could reflect other features such as deterministic transitions or negative rewards rather than a cost-benefit calculation, as the authors note. The paper does not measure effort or cognitive cost directly, only reward. Participants were online workers with limited trials, and model comparison for the original task partly favoured a pure model-free model.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • The classic task gives little incentive to plan.

    In the original two-step task, planning barely earns more reward: simulated model-free and model-based agents were rewarded on about 0.558 and 0.557 of trials, and more model-based control did not raise reward across nearly the whole parameter range of the original task and two variants.

    Evidence for the claim as stated.

  • People may use planning more when it is worth the effort.

    When a redesigned task made planning pay, people planned more: model-based weight correlated with reward in the new task (r = 0.55) but not the original (r = 0.10), and median weight was higher (0.48 vs 0.27).

    Evidence for the claim as stated.

  • People may use planning more when it is worth the effort.

    When a redesigned task made planning pay, people planned more: model-based weight correlated with reward in the new task (r = 0.55) but not the original (r = 0.10), and median weight was higher (0.48 vs 0.27).

    Scope note — Between-group correlational comparison of online workers; other task features could explain the difference, and effort was not measured.

    Limits the claim's scope: a different population, assay, or outcome.

Discoveries this paper informs or conflicts with

Related papers in this topic

Same topic cluster — not a recommendation engine.