Skip to content
PaperFren

How can we get better at finding a stranger on CCTV?

Open paper intelligence

Finding an unfamiliar person in busy CCTV footage is very error-prone, but giving searchers several different photos of the target and using higher-resolution video both help.

Source

Face search in CCTV surveillance

Mileva M, Burton AM · Cognitive research: principles and implications · 2019

doi.org/10.1186/s41235-019-0193-0Read the full paper ↗8 citationscc by

Study at a glance

Design
Human experiment — Four lab experiments manipulating the search-target display (1 vs 3 photos; 1 vs 16 photos; still vs video target with SD vs HD CCTV; wanted/missing poster context vs none) with 50% target-present trials in 2-minute CCTV clips
N
Separate samples per study: 50 in Study 1, 24 in Study 2, 40 in Study 3 (half saw SD and half HD footage), 24 in Study 4; mostly university students at York
Population
University students (and a few staff) with no prior familiarity with the target volunteers
Outcome
Accuracy at identifying the target in target-present clips and correctly rejecting target-absent clips; crowd accuracy from aggregated majority votes

Structured fields used in claim comparison tables when every cited study has a complete layer.

What they did

Students searched 2-minute greyscale clips of real CCTV from a busy rail station for a target person who was present on half the trials, with full control to pause and rewind. Across four studies the researchers varied what they saw of the target: one versus three ID photos, one versus sixteen varied photos, a still versus a short head-turning video (with standard- or high-definition CCTV), and photos with or without a 'wanted' or 'missing person' back-story. They also pooled responses across simulated groups to test a wisdom-of-crowds effect.

What they found

Performance was poor overall; with one photo, people correctly said the target was absent only 57% of the time. Three photos improved both hits and correct rejections, and sixteen photos gave a similar gain of about 10% but no more than three. A moving target video gave no benefit over a still, while high-definition CCTV clearly beat standard definition. Wanted or missing context did not help even though participants remembered it. Majority votes from groups improved accuracy, except that crowds added little with standard-definition footage.

The limits

What it doesn't show

The comparison between 3 and 16 photos came from two different experiments with different samples and stimuli, so the study cannot establish that more photos add nothing. Samples were small student groups (24 to 50 per study) rather than trained CCTV operators, and the natural footage varied uncontrollably in crowding and lighting. Target prevalence was 50%, far higher than in real surveillance, which could change how often people report a match. The finding that video targets did not help may reflect the load of watching two moving displays rather than a lack of useful motion information.

Key terms

Within-person variability
The natural differences in how the same person looks across photos, lighting, poses and occasions.
Target-present / target-absent trial
Trials where the sought person does or does not appear; absent trials test whether people wrongly pick someone.
Misidentification
Choosing the wrong person as the target, a serious error in forensic settings.
Wisdom of the crowds
Pooling many people's independent judgements, for example by majority vote, to get more accurate decisions than individuals.
Bayes factor
A ratio expressing how much more the data support one hypothesis than another; here used to show evidence for no effect.

Flashcards

1 / 11

0 of 11 answers reviewed

Research intelligence for this paper

See its role on concept claims, tensions it is part of, placement history, and related discoveries.

Open paper intelligence

Quiz yourself

1 / 5

In Study 1, what happened when searchers had three photos of the target rather than one?

Common questions

Why would several photos help if you are still looking for one person?

Different photos show how the target's appearance varies, which may help build a more general representation of their face, similar to how we become familiar with people.

Isn't a video just lots of photos?

A video comes from one occasion, so it lacks the variation in lighting, hairstyle and time that separate photos give; searchers also tended to freeze the target video, suggesting two moving displays were too demanding.

Why did the crowd method barely help with low-resolution footage?

The authors suggest the identity information simply was not present in standard-definition video, so pooling more viewers could not recover it.

More on Face processing