Which brain signal tracks how surprising a word's meaning is?
Readers slowed down on words that made a mini-story implausible whether or not the word was related to the context, and this pattern matched the P600 brain response rather than the N400, as the authors' model predicted.
Source
Neurobehavioral Correlates of Surprisal in Language Comprehension: A Neurocomputational Model
Study at a glance
- Design
- Computational / modelling — Recurrent neural-network model producing N400, P600 and Surprisal estimates, compared with a previously published ERP study (re-analysed with regression ERPs) and a new self-paced reading replication using the same German materials
- N
- N=31 · 31 native German speakers in the new self-paced reading experiment; the ERP data come from an earlier study (Delogu et al.) and the model was trained in 10 instances
- Population
- Native German-speaking students at Saarland University (reading experiment); simulated comprehender (model)
- Outcome
- Word-by-word reading times on the target and following word; plausibility judgements; model N400, P600 and Surprisal estimates
Structured fields used in claim comparison tables when every cited study has a complete layer.
What they did
The authors built a neural-network model of word-by-word comprehension in which a retrieval stage (linked to the N400 brain response) looks up a word's meaning and an integration stage (linked to the P600) adds it to the unfolding interpretation; the model also computes how surprising each updated interpretation is. They generated predictions for an earlier German ERP study with three kinds of two-sentence stories: a plausible baseline, an implausible version where the target word was still associated with the context, and an implausible version where it was unrelated. They re-examined that ERP data with regression-based ERPs, and ran a new self-paced reading version of the same study with 31 participants.
What they found
The model predicted an N400 effect only for the unrelated condition (driven by word association) but a P600 effect and higher Surprisal for both implausible conditions (driven by plausibility). Once overlap between the two ERP components was separated statistically, the ERP data matched this pattern. In the reading experiment, participants judged 91% of baseline stories plausible versus 24% and 8% of the two implausible versions, and they read both implausible target words more slowly than the baseline; plausibility predicted reading time while association did not.
The limits
What it doesn't show
The model was trained on a tiny artificial world of 16 words and 160 sentences, so it demonstrates a principle rather than a realistic model of language. The link between Surprisal and the P600 is qualitative: the design had only three conditions and cannot show that the P600 grows smoothly with graded surprise, which the authors leave open. The ERP evidence depends on a regression method to undo overlap between the N400 and P600 in the raw data, and the reading-time and ERP data came from different people, so the two measures were never recorded together.
Key terms
- N400
- A negative-going ERP component around 400 ms after a word; in this account it reflects how easily the word's meaning is retrieved given the context.
- P600
- A later positive-going ERP component around 600 ms and beyond; here it is taken to reflect the effort of integrating a word into the interpretation.
- Surprisal
- How unexpected something is, measured as the negative log probability; here applied to how improbable the updated interpretation is given what came before.
- Retrieval-Integration account
- A theory that each word is processed in a cycle of retrieving its meaning (N400) and then integrating it into the utterance meaning (P600).
- Self-paced reading
- A method in which readers press a key to reveal each next word, so the time spent on each word indexes processing difficulty.
- Spatiotemporal component overlap
- The problem that different ERP components happening at similar times and scalp locations sum together, so one can mask the other in the recorded signal.
Flashcards
0 of 10 answers reviewed
Research intelligence for this paper
See its role on concept claims, tensions it is part of, placement history, and related discoveries.
Quiz yourself
According to the Retrieval-Integration account, the N400 mainly reflects:
Common questions
If the unrelated condition didn't show a clear P600 in the raw ERPs, why do the authors say it had one?
Because a large, sustained N400 negativity overlapped the P600 time window and cancelled it out. When regression-based ERPs isolated the contribution of plausibility alone, a P600 effect appeared for that condition too.
Doesn't earlier research link Surprisal to the N400?
Yes, but the authors argue the N400 only correlates with Surprisal because expected words are usually easier to retrieve. In this design, association and plausibility were pulled apart, and only the P600 and reading times followed plausibility.
Why is the model 'comprehension-centric'?
Its Surprisal is computed from changes in meaning representations that encode world knowledge, not from word-sequence probabilities like a typical language model.
More on Language