How do we tell our own voice apart from others?
We recognize our own voice better than others' voices when low-frequency pitch and speech details are removed, leaving only the highest vocal frequencies.
Source
Acoustic cues for the recognition of self-voice and other-voice
What they did
To investigate self-voice recognition, researchers recruited 30 participants who were divided into familiar groups. The researchers recorded each person reading short words and then created 15 altered versions of these clips by shifting the pitch by up to 4 semitones and applying low-pass, high-pass, or normal frequency filters. Participants performed 300 test trials where they listened to these degraded clips and had to quickly identify whether the voice belonged to themselves or one of their 4 colleagues.
What they found
Overall, filtering out frequencies and shifting pitch made voice recognition more difficult. However, when low-frequency speech components were completely removed, pitch shifts had no impact on accuracy, and participants recognized their own voice significantly better than their colleagues' voices (60.17% versus 49.67% accuracy). In contrast, when normal or low-frequency bands were intact, there was no significant performance difference between recognizing oneself versus others.
The limits
What it doesn't show
This study does not show how biological sex might affect voice recognition because the sample of 30 participants was not analyzed for sex differences. Additionally, because participants listened to recorded playbacks, the study cannot represent how we recognize our own voice in daily life when we hear it live through bone conduction. Finally, the design cannot rule out that the advantage in recognizing oneself over colleagues is driven entirely by auditory familiarity rather than a dedicated cognitive mechanism for self-recognition.
Key terms
- Self-voice recognition
- The cognitive capacity to identify one's own voice as belonging to oneself.
- Fundamental frequency
- The rate of vocal fold vibration that determines the perceived pitch of a voice.
- Formants
- Resonant frequency bands of the vocal tract that shape vowels and determine vocal timbre.
- Mora
- A phonological unit of sound in languages like Japanese that determines syllable weight and timing.
- Bone conduction
- The transmission of sound waves through the bones of the skull directly to the inner ear, which influences how a person perceives their own speaking voice.
Flashcards
Want these cards to stick?
Save the deck to NoteFren and study it with spaced repetition.
Quiz yourself
Which brain region has been consistently activated during both self-face and self-voice recognition, suggesting a role in abstract multimodal self-representation?
Common questions
Why did high-frequency filtering make it easier to recognize one's own voice than a colleague's?
Vocal frequencies above 2500 Hz are linked to the unique, stable anatomy of the larynx. Listeners have immense familiarity with their own voice, which helps them recognize these stable features even when other speech cues are missing.
How does altering the pitch affect how we recognize voices?
In normal and low-frequency recordings, shifting the pitch by up to 4 semitones significantly decreases recognition accuracy. However, in high-frequency recordings, changing the pitch has no effect on recognition.
How did researchers prevent participants from visually identifying the speakers?
Participants had to keep their eyes closed during the test trials and identify the speaker purely based on brief 600-millisecond audio recordings.
Why did the researchers use voices of familiar colleagues rather than strangers?
Using colleagues who knew each other's voices ensured that any differences in recognition were due to the unique properties of self-voice processing rather than simply comparing a highly familiar self-voice to unfamiliar other-voices.
More on Perception