Neural network models of the brain
Which vision models represent objects like the brain's IT cortex?
Open access · cc by · source: Europe PMC
Only a deep network trained with a million labelled images came close to the category structure of inferior temporal cortex; unsupervised and hand-engineered models did not.
Study at a glance
- Design
- Computational / modelling — Representational similarity analysis comparing model RDMs with existing human fMRI and monkey cell-recording RDMs for 96 object images, plus cross-validated feature reweighting
- N
- 37 model representations were compared on a set of 96 images. The brain data were reused from earlier human fMRI and monkey recording studies, and the monkey data came from two animals.
- Population
- Computational vision models, compared with human IT (fMRI) and monkey IT (cell recordings)
- Outcome
- Kendall tau-a correlation between model and IT representational dissimilarity matrices, and a categoricality index
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
The not-strongly-supervised models all correlated only weakly with IT, and none showed IT's strong animate/inanimate split or its face cluster spanning humans and animals. Human IT had a categoricality index of 0.4, while all these models scored below 0.16. Layer 7 of the deep supervised network matched human IT better (tau-a 0.24) than the best combination of other models (0.17). Reweighting the deep network's layers together with category-discriminant features produced a model (tau-a 0.38) that reached the noise ceiling for human IT.
Methodology
The authors tested 37 model representations, including neuroscience-inspired models like HMAX, classic computer-vision features like SIFT and GIST, and every layer of a deep convolutional network trained on ImageNet. They used representational similarity analysis, comparing how each model and the brain separate the same 96 object images, against human IT fMRI and monkey IT recordings. They also tried reweighting and remixing model features with cross-validation to see whether a better IT match could be built.
Limitations
The stimulus set is small (96 images of isolated objects) and was split equally between animate and inanimate objects, which shapes the categoricality measure. Only one deep supervised network was tested, so supervision is confounded with that architecture and its large training set. The best-fitting model was partly fitted to the IT data, though with cross-validation. The noise ceiling could not be estimated for monkey IT because data came from only two animals.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Category-trained networks capture IT's structure; unsupervised ones did not.
Supervised training on object categories matters for higher visual cortex: unsupervised models correlated weakly with IT and none showed IT's animate/inanimate split, while the deep supervised network matched human IT better, and a reweighted version reached the noise ceiling.
Evidence for the claim as stated.
Supervision is essential for IT but not obviously for V1: the IT study found only the supervised network matched, while in V1 a network trained only on neural data (no object labels) came close to VGG (49.8% vs 51.6%) though it needed more data.
Evidence for the claim as stated.
Open questions
Tensions this paper is part of
From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.
Supervision is essential for IT but not obviously for V1: the IT study found only the supervised network matched, while in V1 a network trained only on neural data (no object labels) came close to VGG (49.8% vs 51.6%) though it needed more data.
Discoveries this paper informs or conflicts with
- Deep networks trained on object labels predict visual cortex better than hand-built and unsupervised models
This paper informs this development.
Related papers in this topic
Same topic cluster — not a recommendation engine.