Skip to content
Marius Comper

Acoustic Physics of the Vocal Tract

Two Notes in the Mouth

When you voice a vowel, your vocal folds produce nothing more than a raw, buzzing saw wave, much like the reed of a saxophone. What turns that buzz into “a”, “e”, or “u” is the physical division of your vocal tract into two resonant chambers: one behind the tongue and one in front of it. Each chamber amplifies its own resonant frequency, known as a formant. Together, these two notes form an acoustic chord that the human brain instantly decodes as speech.

3.01 octaves separating the two dominant frequencies in “i”, the widest acoustic spread
1,530 Hz the leap in the second frequency when gliding from “u” to “i”, without moving the jaw
2 distinct central vowels (“ă” and “î”) anchoring the median axis of Romanian
+16% higher formant frequencies in female speech due to a 3 cm shorter vocal tract

How the Tongue Sculptures Sound

The human pharynx and mouth form an acoustic tube approximately 17.5 centimetres long in adult males and 14.5 centimetres in females. When air from the lungs passes through the larynx, the vocal folds vibrate at a low fundamental frequency (around 115 Hz for men and 210 Hz for women). This vibration is rich in higher harmonics, but without the mouth cavity it sounds like a mechanical drone.

By raising the tongue or opening the jaw, we divide the tube into a dual Helmholtz acoustic resonator:

The First Formant (F₁) is determined by jaw opening and the pharyngeal space behind the tongue. When the mouth opens wide, as in “a”, F₁ climbs to around 780 Hz. When the mouth closes tightly, as in “i” or “u”, F₁ drops to 280–310 Hz, deep in the bass register.

The Second Formant (F₂) is determined by the front-to-back position of the tongue. When the tongue arches forward toward the teeth for “i”, the oral chamber becomes very small and rings at a high 2,250 Hz. When the tongue retracts toward the throat for “u”, the forward chamber expands, causing F₂ to drop to just 720 Hz.

Vocal Resonance Laboratory

Drag the probe across the acoustic plane to hear real-time formant synthesis or track the tongue trajectory across any word.

Click & drag anywhere on the acoustic plane

Track the Tongue Across Words

Choose a Romanian word or enter your own text to trace the physical path of the tongue moving inside the vocal cavity.

Vowels traversed: 4 Total acoustic distance: 3,412 cents (≈ 34.1 semitones) What are cents? The logarithmic unit of pitch intervals: 100 cents = 1 semitone (one piano key), and 1,200 cents = 1 octave. This number measures the acoustic distance mouth resonances travel between vowels.

The Central Spine: Why Romanian Is Unusual

In most Romance languages, the vowel system forms a hollow perimeter. Spanish and Italian push all vowels to the outer edges: two front vowels (/i/, /e/), two back vowels (/u/, /o/), and a single open vowel at the base (/a/). The center of the mouth remains completely unused.

Romanian is unique among its family in having developed two distinct central vowels placed directly along the neutral midline of the mouth:

The vowel “ă” (/ə/) occupies the exact acoustic center of the mouth cavity (F₁ ≈ 530 Hz, F₂ ≈ 1,410 Hz). The tongue rests in neutral equilibrium, displaced neither forward nor backward.

The vowel “î” or “â” (/ɨ/) raises the tongue toward the palate while maintaining this central alignment (F₁ ≈ 340 Hz, F₂ ≈ 1,510 Hz). This divides the mouth uniquely: F₁ drops almost as low as in “i”, but F₂ stays suspended midway between front and back resonances.

Vowel Reference word F₁ (Male) F₂ (Male) F₁ (Female) F₂ (Female) Ratio F₂/F₁ Interval
i fir 280 Hz 2,250 Hz 320 Hz 2,750 Hz 8.04× 3.01 octaves
e tren 450 Hz 1,900 Hz 530 Hz 2,300 Hz 4.22× 2.08 octaves
â / î cânt 340 Hz 1,510 Hz 390 Hz 1,720 Hz 4.44× 2.15 octaves
ă casă 530 Hz 1,410 Hz 600 Hz 1,620 Hz 2.66× 1.41 octaves
a pas 780 Hz 1,320 Hz 900 Hz 1,550 Hz 1.69× 0.76 octaves
o nord 480 Hz 920 Hz 550 Hz 1,050 Hz 1.92× 0.94 octaves
u nor 310 Hz 720 Hz 360 Hz 820 Hz 2.32× 1.22 octaves

Why Vowels Are Continuous Coordinates

The alphabet treats vowels as discrete boxes. In the physical reality of human speech, vowels form a continuous two-dimensional surface. When gliding from “i” to “u”, you do not switch channels: the tongue slides smoothly from teeth to throat, sweeping the second formant continuously from 2,250 Hz down to 720 Hz.

Using the formant isolation buttons above, you can hear that the first formant (F₁) acts as a muffled bass tone, while the second formant (F₂) acts as a variable whistle pitch. The human ear and brain merge these two frequencies in milliseconds to decode speech.

Method and Sources

Formant values for standard Romanian are based on published experimental phonetics and acoustic studies by Laurenția Dascălu-Jinga (Tratat de Fonetică a Limbii Române, Romanian Academy), Andrei Avram (Cercetări asupra sistemului fonologic al limbii române), Margaret E. L. Renwick (Vowels of Romanian: Historical Phonology and Experimental Phonetics, Cornell University, 2012), and Ioana Chițoran (The Phonology of Romanian, Mouton de Gruyter, 2001).

The acoustic synthesis engine implements Gunnar Fant’s Source-Filter model (1960) using digital biquad bandpass filters with Q factors calibrated to the physiological bandwidth of the human vocal tract. Acoustic intervals are computed via base-2 logarithms: \(\text{cents} = 1200 \times \log_2(F_2 / F_1)\).