Skip to content
PaperFren

Concept

Robot learning from rewards, demonstrations and humans

9 studiesEvidence last moved Sep 27, 2026

Robot learning covers how robots acquire skills from reward signals, human demonstrations, human ratings or even evolutionary inheritance rather than hand-coded control. This page also covers robot motor learning; the evidence is mostly simulation studies plus small real-robot demonstrations and user studies with 8–26 sessions.

Students often picture robots learning from a clean reward. In practice, rewards are shaped, humans provide noisy signals, and simulation results don't always transfer, which these studies show concretely.

Studies

9

Findings

5

10 supporting · 0 challenging · 0 qualifying citations

Open tensions

1

Latest change

Concept page published

Robot learning from rewards, demonstrations and humans

Currently

What we know

  1. What you reward, and what the robot senses, shapes what it learns.
  2. Negative reward is not a free safety signal.
  3. Copying demonstrations alone breaks when the scene changes.
  4. People can be the reward, but how they give it decides the outcome.
  5. Curiosity, inheritance and few-shot examples are viable but lightly tested.

Largest unresolved question

Simulation rankings do not fully carry to real robots: the full Coulomb-plus-vision-plus-proximity model was best in simulation, but on the real robot the model without the proximity penalty did best.

Common misconceptions

  • Results in simulation show a method works on robots.

    Most studies here (EPG, navigation, Lamarckian evolution, nociception) were simulation-only, and the one sim-to-real comparison found the best model changed.

  • A non-significant difference between human ratings and a cost function proves they're equivalent.

    The user study's authors note small groups; lack of a significant difference is not proof of equivalence.

Related