Is amygdala activity to emotional faces stable enough to be a biomarker?
Although emotional faces reliably activated the amygdala and deactivated the subgenual cingulate on average, how strongly each individual responded was mostly not consistent from one scan to the next.
Source
Unreliability of putative fMRI biomarkers during emotional face processing
Study at a glance
- Design
- Other — Test-retest reliability study: the same healthy adults did three emotional-face fMRI tasks twice per session on two days about two weeks apart; intraclass correlations computed per region.
- N
- N=29 · 29 participants analysed after exclusions from 35 recruited; per-task analyses used 27 (emotion identification), 29 (emotion matching) and 28 (gender classification).
- Population
- Healthy right-handed adults aged 18 to 40 with no psychiatric history, recruited in London
- Outcome
- Intraclass correlation (ICC) of BOLD responses in the left and right amygdala and subgenual anterior cingulate, with the right fusiform face area as a control
Structured fields used in claim comparison tables when every cited study has a complete layer.
What they did
Twenty-nine healthy young adults were scanned on two days roughly two weeks apart, doing three common face tasks twice each day: naming the emotion, matching emotional faces versus shapes, and judging the gender of emotional faces. The researchers extracted each person's response in the left and right amygdala and the subgenual anterior cingulate (sgACC), using both group-peak and anatomical regions, and computed intraclass correlations between days and between runs. The right fusiform face area served as a positive control region.
What they found
All three tasks produced the expected group-level amygdala activation and sgACC deactivation, but between-day reliability was poor in most cases, with seven of nine functional-region estimates well below the 0.4 threshold for moderate reliability. Reliability was also mostly poor within a single session, even between runs minutes apart. By contrast, the fusiform face area was at least moderately reliable in 13 of 15 tests and excellent in both block-design tasks. The only between-day result that was moderately reliable under both region-definition methods was the sgACC in the gender task, and even that never reached excellent reliability.
The limits
What it doesn't show
The sample was healthy young adults, so reliability in depressed patients, the group where biomarkers matter, is untested and might differ. Only three regions and three face tasks were examined, and a scanner at 1.5 T with restricted coverage was used, so other regions, tasks, field strengths or measures such as habituation might be more stable. Changes in mood or anxiety between sessions and habituation to the faces could partly explain the instability, and the study was powered only to detect a reliability of about 0.5.
Key terms
- Test-retest reliability
- How consistently a measure ranks the same individuals when it is repeated.
- Intraclass correlation (ICC)
- A statistic from 0 to 1 describing how much of the variation in repeated measurements reflects stable differences between people rather than noise; below 0.4 is conventionally poor.
- Biomarker
- A measurable biological signal used to predict or track a clinical outcome such as treatment response.
- Subgenual anterior cingulate cortex (sgACC)
- A region below the front of the corpus callosum linked to mood and repeatedly proposed as a predictor of antidepressant response.
- Region of interest (ROI)
- A predefined brain area from which activity values are extracted, defined either from group activation peaks or from anatomy.
Flashcards
0 of 10 answers reviewed
Research intelligence for this paper
See its role on concept claims, tensions it is part of, placement history, and related discoveries.
Quiz yourself
What conventional ICC value marks the lower bound of moderate reliability?
Common questions
How can a region activate reliably across the group but be unreliable in individuals?
A group average can be robust even when each person's response bounces around between sessions; reliability asks whether people keep the same rank order, which is a different question.
Why does unreliability matter for predicting treatment response?
If a person's measurement changes substantially from day to day, it cannot accurately predict how that particular person will respond to treatment.
Why include the fusiform face area?
As a positive control: its high reliability shows the scanner, tasks and analysis were capable of producing stable measurements, so the poor amygdala results are not simply a broken setup.
More on Emotion circuits