How can we get better at finding a stranger on CCTV?
Finding an unfamiliar person in busy CCTV footage is very error-prone, but giving searchers several different photos of the target and using higher-resolution video both help.
Source
Face search in CCTV surveillance
Study at a glance
- Design
- Human experiment — Four lab experiments manipulating the search-target display (1 vs 3 photos; 1 vs 16 photos; still vs video target with SD vs HD CCTV; wanted/missing poster context vs none) with 50% target-present trials in 2-minute CCTV clips
- N
- Separate samples per study: 50 in Study 1, 24 in Study 2, 40 in Study 3 (half saw SD and half HD footage), 24 in Study 4; mostly university students at York
- Population
- University students (and a few staff) with no prior familiarity with the target volunteers
- Outcome
- Accuracy at identifying the target in target-present clips and correctly rejecting target-absent clips; crowd accuracy from aggregated majority votes
Structured fields used in claim comparison tables when every cited study has a complete layer.
What they did
Students searched 2-minute greyscale clips of real CCTV from a busy rail station for a target person who was present on half the trials, with full control to pause and rewind. Across four studies the researchers varied what they saw of the target: one versus three ID photos, one versus sixteen varied photos, a still versus a short head-turning video (with standard- or high-definition CCTV), and photos with or without a 'wanted' or 'missing person' back-story. They also pooled responses across simulated groups to test a wisdom-of-crowds effect.
What they found
Performance was poor overall; with one photo, people correctly said the target was absent only 57% of the time. Three photos improved both hits and correct rejections, and sixteen photos gave a similar gain of about 10% but no more than three. A moving target video gave no benefit over a still, while high-definition CCTV clearly beat standard definition. Wanted or missing context did not help even though participants remembered it. Majority votes from groups improved accuracy, except that crowds added little with standard-definition footage.
The limits
What it doesn't show
The comparison between 3 and 16 photos came from two different experiments with different samples and stimuli, so the study cannot establish that more photos add nothing. Samples were small student groups (24 to 50 per study) rather than trained CCTV operators, and the natural footage varied uncontrollably in crowding and lighting. Target prevalence was 50%, far higher than in real surveillance, which could change how often people report a match. The finding that video targets did not help may reflect the load of watching two moving displays rather than a lack of useful motion information.
Key terms
- Within-person variability
- The natural differences in how the same person looks across photos, lighting, poses and occasions.
- Target-present / target-absent trial
- Trials where the sought person does or does not appear; absent trials test whether people wrongly pick someone.
- Misidentification
- Choosing the wrong person as the target, a serious error in forensic settings.
- Wisdom of the crowds
- Pooling many people's independent judgements, for example by majority vote, to get more accurate decisions than individuals.
- Bayes factor
- A ratio expressing how much more the data support one hypothesis than another; here used to show evidence for no effect.
Flashcards
0 of 11 answers reviewed
Research intelligence for this paper
See its role on concept claims, tensions it is part of, placement history, and related discoveries.
Quiz yourself
In Study 1, what happened when searchers had three photos of the target rather than one?
Common questions
Why would several photos help if you are still looking for one person?
Different photos show how the target's appearance varies, which may help build a more general representation of their face, similar to how we become familiar with people.
Isn't a video just lots of photos?
A video comes from one occasion, so it lacks the variation in lighting, hairstyle and time that separate photos give; searchers also tended to freeze the target video, suggesting two moving displays were too demanding.
Why did the crowd method barely help with low-resolution footage?
The authors suggest the identity information simply was not present in standard-definition video, so pooling more viewers could not recover it.
More on Face processing