Skip to content
PaperFren

Robot learning

Can ordinary people teach a robot a skill just by rating its tries?

Vollmer AL, Hemion NJ · Frontiers in robotics and AI · 2018

Open access · cc by · source: Europe PMC

Non-expert users giving simple one-to-five star ratings taught a robot the ball-in-cup trick about as well as a carefully engineered camera-based cost function.

Study at a glance

Design
Human experiment — Humanoid robot Pepper optimised a ball-in-cup movement using either naive users' star ratings as the cost or a tuned camera-based cost function; success measures compared between setups and across users' rating strategies.
N
N=26 · 26 non-expert participants each taught one session of 82 rated movements; compared with 15 camera-optimised sessions.
Population
Adult non-experts aged 19-70 recruited around Bielefeld University; Pepper humanoid robot.
Outcome
Final hit/miss and ball-cup distance, number of hits, roll-outs until first hit; correlation of user ratings with camera-measured distance.

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

None of the five success measures differed significantly between human ratings and the camera cost function. Of 26 user sessions, 17 converged successfully, 6 converged prematurely and 3 failed; the camera condition also had 2 failures out of 15. Most participants rated by how close the ball came to the cup, and all of them were in successful sessions, while all participants who rated each attempt relative to the previous one produced premature convergence. Ratings correlated with camera-measured distance at about 0.72 on average, and this correlation was significantly higher in successful than in failed sessions.

Methodology

A Pepper humanoid robot learned the ball-in-cup game using a standard movement representation (dynamic movement primitives) optimised by a simple evolution strategy (CMA-ES), starting from an imperfect demonstrated throw. In one condition a tuned two-camera system measured ball-to-cup distance as the cost; in the other, 26 naive participants rated each of the robot's attempts from one to five stars on its tablet, with no instructions about how to rate. The authors compared five success measures between setups and related participants' self-reported rating strategies and rating accuracy to learning success.

Limitations

Only one task on one robot was tested, so it is unknown whether rating-based learning works for other skills. The comparison is a non-significant difference with small groups, which is not proof that the two setups are equivalent. Users were in a lab with an experimenter present, not at home, and several found the near-identical attempts within a batch confusing or gave identical scores that carry no learning signal. The strategy-outcome links rest on self-report and very small subgroups.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • People can be the reward, but how they give it decides the outcome.

    Human feedback can substitute for an engineered cost function but its quality matters: human ratings produced success rates not significantly different from a camera cost function, but raters who scored relative to the previous attempt caused premature convergence; with EEG-based feedback, robot errors correlated strongly with decoder misclassifications.

    Evidence for the claim as stated.

Related papers in this topic

Same topic cluster — not a recommendation engine.