Face processing
How can we get better at finding a stranger on CCTV?
Open access · cc by · source: Europe PMC
Finding an unfamiliar person in busy CCTV footage is very error-prone, but giving searchers several different photos of the target and using higher-resolution video both help.
Study at a glance
- Design
- Human experiment — Four lab experiments manipulating the search-target display (1 vs 3 photos; 1 vs 16 photos; still vs video target with SD vs HD CCTV; wanted/missing poster context vs none) with 50% target-present trials in 2-minute CCTV clips
- N
- Separate samples per study: 50 in Study 1, 24 in Study 2, 40 in Study 3 (half saw SD and half HD footage), 24 in Study 4; mostly university students at York
- Population
- University students (and a few staff) with no prior familiarity with the target volunteers
- Outcome
- Accuracy at identifying the target in target-present clips and correctly rejecting target-absent clips; crowd accuracy from aggregated majority votes
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
Performance was poor overall; with one photo, people correctly said the target was absent only 57% of the time. Three photos improved both hits and correct rejections, and sixteen photos gave a similar gain of about 10% but no more than three. A moving target video gave no benefit over a still, while high-definition CCTV clearly beat standard definition. Wanted or missing context did not help even though participants remembered it. Majority votes from groups improved accuracy, except that crowds added little with standard-definition footage.
Methodology
Students searched 2-minute greyscale clips of real CCTV from a busy rail station for a target person who was present on half the trials, with full control to pause and rewind. Across four studies the researchers varied what they saw of the target: one versus three ID photos, one versus sixteen varied photos, a still versus a short head-turning video (with standard- or high-definition CCTV), and photos with or without a 'wanted' or 'missing person' back-story. They also pooled responses across simulated groups to test a wisdom-of-crowds effect.
Limitations
The comparison between 3 and 16 photos came from two different experiments with different samples and stimuli, so the study cannot establish that more photos add nothing. Samples were small student groups (24 to 50 per study) rather than trained CCTV operators, and the natural footage varied uncontrollably in crowding and lighting. Target prevalence was 50%, far higher than in real surveillance, which could change how often people report a match. The finding that video targets did not help may reflect the load of watching two moving displays rather than a lack of useful motion information.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Not yet placed on a claim. This paper has study layers, but no concept page yet cites it as support, challenge, or qualifier.
Related papers in this topic
Same topic cluster — not a recommendation engine.