Topic
Reinforcement learning research, explained
11 open-access reinforcement learning studies, each with a flashcard deck and a quiz.
- Do reward cues push us to approach rather than just to act?
Cues that predicted money made people more likely to approach but less likely to withdraw, so Pavlovian cues bias specific kinds of action rather than simply energising behaviour.
- When is it worth stopping to think before acting?
A model in which the brain only plans when the expected gain from better information beats the reward lost by waiting explains why well-practised actions become fast, inflexible habits.
- Can a predictive map explain planning with simple learning?
Learning values on top of a map of which states tend to follow which lets an agent show flexible, planning-like behaviour using the same prediction-error rule linked to dopamine, and adding offline replay closes the remaining gaps.
- Does planning ahead actually earn more reward in lab tasks?
In the most widely used planning-versus-habit task, planning earns essentially no extra reward, but a redesigned task makes planning pay off and people plan more in it.
- Are habits chunked action sequences run by a goal-directed boss?
People's habitual choices were better explained as whole pre-packaged action sequences chosen by a goal-directed system than as separate single actions valued by a model-free habit system.
- How do people learn whether an adviser is trying to help them?
People's choices were best explained by a learning model that tracks not only how accurate an adviser is but also how quickly the adviser's intentions are changing.
- Do people learn more from outcomes that confirm their choice?
People learned more from news that their choice was right, whether it came from the option they picked or the one they skipped, and this bias made them slower to adapt when rewards switched.
- Can habits look like planning in the two-step task?
Agents that never plan can still produce the behavioural pattern usually taken as proof of planning, so the standard test for 'model-based' choice can be fooled.
- Do teenagers learn from rewards and punishments like adults?
Adolescents' learning was best explained by a simple reward-tracking algorithm, while adults also learned from outcomes they didn't choose and from context, which helped them avoid punishments.
- Can curiosity drive a humanoid robot to learn how its body moves?
Rewarding a real humanoid robot for improving its own model of how its movements turn out made it explore more efficiently than random or least-tried exploration, and led it to discover and probe a table on its own.
- Do pain signals or punishments help a robot learn to reach?
Giving a learning robot arm pain-like sensor inputs made it more accurate and safer, while subtracting punishment from its reward made learning worse.