Neural network models of the brain
Do deep neural networks predict early visual neurons best?
Open access · cc by · source: Europe PMC
Deep convolutional networks predicted monkey V1 responses to natural images far better than classic Gabor-style models, which suggests V1 computation involves several nonlinear steps.
Study at a glance
- Design
- Computational / modelling — Model comparison on held-out images: pretrained VGG-19 features with a regularised GLM readout, CNNs fitted directly to spikes, and LNP and Gabor filter bank baselines
- N
- N=166 · 166 V1 neurons, selected from 262 isolated in 17 sessions, went into the models. The neurons came from two monkeys.
- Population
- V1 neurons in two awake, fixating rhesus macaques viewing natural images and synthesised textures
- Outcome
- Fraction of explainable variance explained (FEV) in spike counts for held-out images
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
The best VGG layer was conv3_1, the fifth of its 16 convolutional layers, and it explained 51.6% of explainable variance. The data-driven CNN reached 49.8%, the Gabor filter bank 45.6% and the LNP only 16.3%. Careful sparsity-based regularisation of the readout was essential. The VGG model kept its performance with only a fifth of the training data, while the data-driven CNN needed the full dataset. The CNNs' gains over the Gabor model did not depend on whether cells were simple or complex, or on how sharply they were tuned.
Methodology
The authors recorded spiking from V1 neurons in two monkeys while the animals viewed rapid sequences of natural images and textures synthesised from them. They then fitted four model types: a linear-nonlinear Poisson (LNP) model, a Gabor filter bank, a GLM readout from each layer of the ImageNet-trained VGG-19 network (the goal-driven approach), and CNNs trained directly on the neural data (the data-driven approach). Each model was scored on images it had not seen in training.
Limitations
The data come from only two monkeys, and responses were limited to brief, mostly feedforward windows after images flashed for 60 ms, so recurrent and feedback processing is not captured. The best models still leave almost half of the explainable variance unexplained. They also do not tell us which nonlinearities (for example divisive normalisation) they capture, because the CNNs are hard to interpret. Better prediction shows that the representation is similar to V1's, not that V1 is trained on object recognition.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
CNN features beat hand-designed filters for V1, but leave half unexplained.
Deep CNN features predict early visual cortex better than classic models: a VGG layer explained 51.6% of explainable variance in macaque V1, a data-driven CNN 49.8%, a Gabor filter bank 45.6% and an LNP model 16.3%.
Evidence for the claim as stated.
Supervision is essential for IT but not obviously for V1: the IT study found only the supervised network matched, while in V1 a network trained only on neural data (no object labels) came close to VGG (49.8% vs 51.6%) though it needed more data.
Evidence for the claim as stated.
Open questions
Tensions this paper is part of
From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.
Supervision is essential for IT but not obviously for V1: the IT study found only the supervised network matched, while in V1 a network trained only on neural data (no object labels) came close to VGG (49.8% vs 51.6%) though it needed more data.
Discoveries this paper informs or conflicts with
- Deep networks trained on object labels predict visual cortex better than hand-built and unsupervised models
This paper informs this development.
Related papers in this topic
Same topic cluster — not a recommendation engine.