Skip to content
PaperFren

Reinforcement learning

Do pain signals or punishments help a robot learn to reach?

Navarro-Guerrero N, Lowe RJ, Wermter S · Frontiers in neurorobotics · 2017

Open access · cc by · source: Europe PMC

Giving a learning robot arm pain-like sensor inputs made it more accurate and safer, while subtracting punishment from its reward made learning worse.

Study at a glance

Design
Computational / modelling — Simulated 2-joint arm reaching task learned with CACLA+var; hyperparameters evolved by a genetic algorithm per condition, then best settings retrained from 10 random initialisations.
N
No single N: four conditions, each with its best hyperparameter set trained 10 times; the genetic search used 32 individuals per generation for 50 generations.
Population
Simulated two-degree-of-freedom robot arm learning inverse kinematics
Outcome
Final positioning error, perceived nociception (potential for joint damage) and number of steps to reach the target

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

The reward-plus-nociception condition reached the lowest positioning error, with less oscillation, and was the only one significantly better than reward only on that measure. It also tended to reduce potential joint damage and steps to target, though those effects were smaller. Adding punishment slowed convergence and, during learning, raised the potential for damage by almost 50%; the genetic search pushed the punishment weight to its smallest allowed value, and nociceptive inputs partly offset punishment's harm.

Methodology

The authors trained a simulated two-joint arm to reach targets using an actor-critic reinforcement learning algorithm (CACLA). They compared four conditions: reward only; reward minus a punishment for joints near their limits; reward only but with extra 'nociceptive' input units signalling when each joint neared its limit; and both. For each condition a genetic algorithm tuned hyperparameters, and the best settings were retrained from 10 random starts and compared statistically.

Limitations

This is one simple simulated task with two joints, so it is unknown whether the effects hold on a real robot or harder problems. Several differences, such as in damage and speed, were small or not reliable, and each condition was judged from just 10 training runs of one tuned configuration. The explanation that extra inputs help by projecting states into a higher-dimensional space, and that punishment hurts by merging with reward into one number, is the authors' hypothesis rather than something tested.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • What you reward, and what the robot senses, shapes what it learns.

    Shaping the reward with extra structured signals speeds and improves learning: a Coulomb-force reward improved success, collisions and goal distance for a navigating manipulator, and adding nociceptive (damage-sensing) inputs to reward gave the lowest positioning error in a simulated arm.

    Evidence for the claim as stated.

  • Negative reward is not a free safety signal.

    Explicit punishment can hurt: in the simulated arm, adding punishment slowed convergence and raised potential damage during learning by almost 50%, and the parameter search pushed the punishment weight to its minimum.

    Evidence for the claim as stated.

Related papers in this topic

Same topic cluster — not a recommendation engine.