Skip to content
PaperFren

Do we recognise 'animal' before 'bird' before 'duck'?

Open paper intelligence

Coarse, blurry image information was enough to tell animals from cars quickly, but telling birds apart from other animals, and ducks from pigeons, needed fine detail and took longer.

Source

Object Categorization in Finer Levels Relies More on Higher Spatial Frequencies and Takes Longer

Ashtiani MN, Kheradpisheh SR, Masquelier T, et al. · Frontiers in psychology · 2017

doi.org/10.3389/fpsyg.2017.01261Read the full paper ↗10 citationscc by

Study at a glance

Design
Human experiment — Two-alternative forced-choice categorisation of briefly flashed, masked object images at three levels (animal vs car, bird vs non-bird animal, duck vs pigeon), with intact or low-, intermediate- or high-spatial-frequency filtered images, with and without phase noise; the same tasks were run on two computational models
N
No single total N is given. For each image-type set of experiments, 40 subjects did the superordinate tasks (10 per task), 40 did the basic-level tasks (10 per task) and 20 did the subordinate task; the text does not say whether the same people took part across the four image-type sets
Population
Adult volunteers tested at the University of Tehran (age and sex not reported in the text)
Outcome
Categorisation accuracy and reaction time by categorisation level, spatial-frequency band and noise level; model accuracy on the same tasks

Structured fields used in claim comparison tables when every cited study has a complete layer.

What they did

Participants saw a grayscale photo for 12.5 ms, followed by a mask, and pressed a key to categorise it at one of three levels: animal versus car (superordinate), bird versus cat or dog (basic), or duck versus pigeon (subordinate). Images were shown intact or filtered to keep only low, intermediate or high spatial frequencies, and in further experiments phase noise was added to avoid ceiling effects. Two computational object-recognition models based on the HMAX architecture were trained and tested on the same tasks to see whether the human pattern follows from the information available in each frequency band.

What they found

For animal-versus-car decisions, accuracy was already high with low-frequency (blurry) images and did not improve with finer detail. For the basic and subordinate levels, accuracy was poor with low frequencies and rose as higher frequencies were available, most steeply for duck versus pigeon. Reaction times rose from low to high frequency bands at every level and were shortest for the superordinate level and longest for the subordinate level. Both models showed accuracy patterns closely resembling the human ones, and adding noise lowered accuracy and slowed responses without changing these patterns.

The limits

What it doesn't show

Only five object classes were used (ducks, pigeons, cats, dogs and cars), and cars were the only non-animal category, so the conclusions may not generalise to other categories or to real-world scenes. The text gives no participant ages, sex or total sample size, and does not say whether the same people did several experiments. The models explain accuracy but not timing, so the claim that superordinate categorisation happens first relies on the separate assumption that low frequencies are processed before high ones. The results conflict with earlier studies favouring a basic-level advantage, which the authors attribute to task differences rather than testing directly.

Key terms

Spatial frequency
How rapidly brightness changes across an image; low frequencies carry coarse shape and layout, high frequencies carry fine edges and texture.
Superordinate, basic and subordinate levels
Levels of category abstraction, for example animal (superordinate), bird (basic) and duck (subordinate).
Entry level of categorisation
The level at which an object is first categorised during visual processing.
Coarse-to-fine processing
The idea that the visual system processes low spatial frequencies before high ones, building a rough sketch before filling in detail.
Backward masking
Presenting a pattern right after a brief stimulus to cut short its processing and limit how long it is available.
HMAX model
A hierarchical computational model of object recognition inspired by the ventral visual stream, alternating filtering and pooling layers.

Flashcards

1 / 11

0 of 11 answers reviewed

Research intelligence for this paper

See its role on concept claims, tensions it is part of, placement history, and related discoveries.

Open paper intelligence

Quiz yourself

1 / 5

Which image information was enough for fast, accurate animal-versus-car decisions?

Common questions

Why flash the images so briefly and then mask them?

Very brief, masked presentation limits processing to the fast feedforward sweep, which makes it easier to see which category level can be reached first and prevents people from simply studying details at leisure.

Why run computer models at all?

The models test whether the human pattern is forced by the information content of each frequency band. Because models with no timing or brain-specific quirks showed the same accuracy pattern, the authors argue the pattern reflects the images' information rather than a peculiarity of human processing.

Why does this challenge the idea of a basic-level advantage?

Classic studies found people name objects fastest at the basic level (bird). Here, without verbal labels, the superordinate level (animal) was fastest, suggesting the basic-level advantage may partly reflect language or semantic processing in those tasks.

More on Perception