Reinforcement learning
When is it worth stopping to think before acting?
Open access · cc by · source: Europe PMC
A model in which the brain only plans when the expected gain from better information beats the reward lost by waiting explains why well-practised actions become fast, inflexible habits.
Study at a glance
- Design
- Computational / modelling — Simulations of an agent combining Kalman model-free learning with model-based value iteration and a cost-benefit arbitrator, run on formalised versions of published animal and human tasks.
- N
- No participants or single dataset; simulated agents are compared qualitatively with published behavioural results.
- Population
- Simulated reinforcement-learning agents modelled on rats and humans.
- Outcome
- Response rates after outcome devaluation, deliberation time, and choice reaction time patterns.
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
After moderate training the simulated agent reduced responding for a devalued outcome, but after extensive training it kept responding, matching the classic finding that overtraining creates habits; deliberation time fell to zero after about 100 training trials. With two equally valued outcomes available concurrently, the agent stayed goal-directed even after long training, as reported in rats. The model also reproduced slower reaction times while searching for a rule than while applying it, and reaction times that grow with the number of options early in training but not later.
Methodology
The authors built an agent with two systems: a fast habitual (model-free) learner that keeps uncertainty estimates for action values, and a slow but accurate goal-directed (model-based) planner. Before each choice, an arbitrator compares the value of perfect information about an action with the cost of deliberating, taken as time multiplied by the average reward rate. They simulated the agent on outcome-devaluation experiments after moderate (40 trials) versus extensive (240 trials) training, a two-lever concurrent schedule, a reversal-learning reaction-time task and choice reaction-time tasks with varying numbers of options.
Limitations
The fits to data are qualitative comparisons with published figures, not quantitative fits to raw data. The model assumes the planner has perfect knowledge of values, which the authors admit fails in tasks like reversal learning, and it predicts a linear rather than the observed logarithmic growth of reaction time with the number of choices. The proposed role of tonic dopamine as the average reward signal is a hypothesis, not tested here.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Habits can be seen as the fast option once planning stops being worth its time.
A speed/accuracy account reproduces the classic overtraining effect: a simulated agent stopped deliberating after about 100 training trials and kept responding for a devalued outcome after extensive training, but stayed goal-directed when two equally valued outcomes were available.
Evidence for the claim as stated.
Discoveries this paper informs or conflicts with
- The standard test of planning versus habit barely rewards planning, and can be passed without it
This paper informs this development.
Related papers in this topic
Same topic cluster — not a recommendation engine.
- Do reward cues push us to approach rather than just to act?
- Can a predictive map explain planning with simple learning?
- Does planning ahead actually earn more reward in lab tasks?
- Are habits chunked action sequences run by a goal-directed boss?
- How do people learn whether an adviser is trying to help them?