Formants: The Shape of a Voice
Someone says a single word from across a crowded room, and you know exactly who it is before you've even turned around. Not because you recognized the sentence. There wasn't one. Just a word, maybe half a word, gone before you could think about it. I keep coming back to what's actually happening in that half second, because it's not really about hearing. It's about resonance.
The chamber, not the source
Your voice starts as a buzz. The vocal cords open and close at some rate, and that rate is pitch, the fundamental. But the buzz on its own doesn't sound like a person, it sounds like a kazoo. What turns it into a voice is everything the buzz passes through on its way out: the throat, the mouth, the space behind the nose. That whole passage is a resonant chamber, and like any chamber, it has a shape, and the shape decides which frequencies get amplified and which ones get swallowed.
Those amplified regions are formants. They're not notes and they're not tones. They're bands of frequency that the chamber happens to favor, purely because of its geometry, the same way a bottle has a pitch it wants to hum at when you blow across the top. Change the size or shape of the chamber even slightly and the formants move. Keep the chamber the same and the formants stay put, no matter what pitch is passing through.
That second part is the whole trick. Formants don't move when your pitch does. Sing a low note, sing a high note, say the same vowel on both, and the formants stay parked in roughly the same place, because your throat didn't get longer and your mouth didn't get bigger just because the pitch went up. The chamber is fixed. The note riding through it isn't.
Pitch is what, formants are how
It helps to keep these two completely separate, because they get blurred together constantly. Pitch is what the vocal cords produce. It's a choice, in the loose sense, you can go up, you can go down, you can slide between notes. Formants are what the vocal tract does to that choice after it's been made. They're not a choice in the same way. They're closer to a fact about your body, fixed by the length of your throat, the size of your mouth, the particular architecture of your sinuses.
Which is why two people can sing the exact same note, checked against a tuner, dead on, and still sound nothing alike. Same pitch. Different chamber. The note is identical and the voice isn't, because the note was never where the identity lived.
A fingerprint made of resonance
This is the part that actually gets me. If formants are set by the physical shape of your vocal tract, and that shape is basically as individual as a fingerprint, then formants are a fingerprint. Not a metaphorical one. An acoustic one, made of resonance instead of ridges, and it's on you every time you make a sound, whether you're whispering or shouting or laughing or half-asleep.
That's why the single word across the room works. You weren't tracking the word's meaning, and you definitely weren't tracking its pitch, since pitch is the least stable thing about how a person talks, it swings with mood and volume and how close they are to you. What stayed constant was the shape of the chamber the word came out of. You recognized the room, not the sentence.
Which raises a question I can't quite let go of. If you stripped away every word someone ever said, every accent, every verbal habit, and left only the bare resonance, the formant structure with nothing riding on it, would you still know who it was? I think you would. I think that's the actual content of "I'd know that voice anywhere," a phrase people say constantly without meaning it literally. They do mean it literally. They're not recognizing vocabulary or cadence or the things a person chooses to say. They're recognizing the shape of the instrument the choosing came out of. Strip the words and the shape is still there, still theirs, still recognizable.
The vowel chart, briefly
There's a standard way to make this visible instead of just asserted. Vowels are plotted on a chart using their first two formants, F1 and F2, and the two together locate a vowel almost like a coordinate. F1 tracks roughly with how open your mouth is, low for a closed vowel like "ee," high for an open one like "ah." F2 tracks roughly with how far forward or back the tongue sits, high for a fronted vowel like "ee," low for a backed one like "oo." Plot enough vowels on those two axes and they cluster into distinct, separated regions, "ah" in one corner, "ee" in another, "oo" somewhere else entirely, each vowel a different resonant shape rather than a different note.
What's interesting is that these regions aren't fixed across every speaker of every language. Different languages, and different dialects within the same language, place their vowels in slightly different spots on that same chart. An "ah" in one accent sits in a marginally different region than the "ah" in another, because the habitual shaping of the vocal tract, where the tongue tends to rest, how open the jaw tends to be, differs by what you grew up hearing and repeating. That's a real part of why an accent doesn't wash out even after decades somewhere else. It's not simply a stubborn memory of pronunciation. It's a chamber that learned a certain shape early and kept it, mostly, resonating the way it first learned to.
Which loops back to the same idea from a different angle. The chamber is close to permanent. What passes through it changes constantly, the words, the pitch, the mood, the meaning, but the shape making the sound recognizable holds steady underneath all of it. That's what a voice actually is, underneath the part anyone can consciously control. Not the note. Not the words. The room the note and the words have to pass through to become someone in particular.