Skip to content
PaperFren

Reinforcement learning

When is it worth stopping to think before acting?

Keramati M, Dezfouli A, Piray P · PLoS computational biology · 2011

Open access · cc by · source: Europe PMC

A model in which the brain only plans when the expected gain from better information beats the reward lost by waiting explains why well-practised actions become fast, inflexible habits.

Study at a glance

Design
Computational / modelling — Simulations of an agent combining Kalman model-free learning with model-based value iteration and a cost-benefit arbitrator, run on formalised versions of published animal and human tasks.
N
No participants or single dataset; simulated agents are compared qualitatively with published behavioural results.
Population
Simulated reinforcement-learning agents modelled on rats and humans.
Outcome
Response rates after outcome devaluation, deliberation time, and choice reaction time patterns.

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

After moderate training the simulated agent reduced responding for a devalued outcome, but after extensive training it kept responding, matching the classic finding that overtraining creates habits; deliberation time fell to zero after about 100 training trials. With two equally valued outcomes available concurrently, the agent stayed goal-directed even after long training, as reported in rats. The model also reproduced slower reaction times while searching for a rule than while applying it, and reaction times that grow with the number of options early in training but not later.

Methodology

The authors built an agent with two systems: a fast habitual (model-free) learner that keeps uncertainty estimates for action values, and a slow but accurate goal-directed (model-based) planner. Before each choice, an arbitrator compares the value of perfect information about an action with the cost of deliberating, taken as time multiplied by the average reward rate. They simulated the agent on outcome-devaluation experiments after moderate (40 trials) versus extensive (240 trials) training, a two-lever concurrent schedule, a reversal-learning reaction-time task and choice reaction-time tasks with varying numbers of options.

Limitations

The fits to data are qualitative comparisons with published figures, not quantitative fits to raw data. The model assumes the planner has perfect knowledge of values, which the authors admit fails in tasks like reversal learning, and it predicts a linear rather than the observed logarithmic growth of reaction time with the number of choices. The proposed role of tonic dopamine as the average reward signal is a hypothesis, not tested here.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • Habits can be seen as the fast option once planning stops being worth its time.

    A speed/accuracy account reproduces the classic overtraining effect: a simulated agent stopped deliberating after about 100 training trials and kept responding for a devalued outcome after extensive training, but stayed goal-directed when two equally valued outcomes were available.

    Evidence for the claim as stated.

Discoveries this paper informs or conflicts with

Related papers in this topic

Same topic cluster — not a recommendation engine.