Frequency Spectrum (0–4000 Hz) — formant peaks
Vowel Chart (F1 vs F2) — IPA cardinal vowels
Vocal Tract Shape (simplified tube model)

About this simulation

Written by MySimulator Team · Reviewed by MySimulator Editorial Review

Last updated: 5 July 2026

This tool synthesises vowel sounds by shaping a glottal source (voiced or whispered) with the resonant peaks — formants F1, F2 and F3 — of a simplified vocal tract tube model. Choosing an IPA vowel preset like /a/, /i/, /u/, /e/ or /o/, or dragging the formant sliders directly, updates the frequency spectrum, the vowel chart position and the tract shape simultaneously.

🔬 What it shows

How three resonant formant frequencies, produced by the shape of the vocal tract, combine with a glottal source to create distinct vowel sounds, visualised as a spectrum, an F1-vs-F2 vowel chart and a tract-shape diagram.

🎮 How to use

Pick a Vowel Preset (/a/ /i/ /u/ /e/ /o/), fine-tune F1, F2, F3, Tract Length and Pitch F0 sliders, choose a Voiced or Whispered Glottal Source, then press ▶ Synthesize Sound or ■ Stop.

💡 Did you know?

Just two formant frequencies, F1 and F2, are enough to distinguish almost every vowel in human speech — that's why the Vowel Chart here plots vowels purely on an F1-vs-F2 grid, the same layout linguists use for the IPA vowel quadrilateral.

Frequently asked questions

What exactly are formants F1, F2 and F3?

They're resonant frequency peaks created by the shape of the vocal tract acting like a tube with varying cross-section — F1 relates mainly to tongue height (jaw openness), F2 to tongue front/back position, and F3 adds finer timbre detail that distinguishes similar vowels.

Why does /i/ have such a low F1 but a very high F2 compared to /a/?

/i/ (as in "see") is produced with a nearly closed jaw and the tongue pushed forward and high, which suppresses F1 (low jaw opening) while pushing F2 up sharply; /a/ (as in "father") does the opposite with an open jaw and low tongue, raising F1 and lowering F2.

What's the difference between the Voiced and Whispered Glottal Source options?

Voiced source simulates the vocal folds vibrating periodically at the Pitch F0 rate, producing a buzzy harmonic source that the tract then filters into recognisable vowel timbre; Whispered source replaces that periodic buzz with broadband noise, which is why whispered vowels sound breathy but far less musical.

How does the Tract Length slider change the sound?

Vocal tract length sets the resonant frequencies of the whole system inversely — a shorter tract (like a child's) shifts all formants higher, while a longer tract (typical of adult males) shifts them lower, which is part of why voices of different ages and sexes sound different even producing the "same" vowel.

Why does changing Pitch F0 not change which vowel you hear?

Pitch F0 is the rate of vocal fold vibration — it determines the fundamental musical note of the voice — while the vowel identity comes from the formant frequencies shaped by the tract. The two are largely independent, which is why you can say the same vowel at any pitch.