Voice Spectrogram Analyzer Online

Reveal the frequency, time, intensity, and estimated formant patterns of two audio samples privately in your browser.

Local only

Your recordings stay on this device. Decoding and analysis run inside the browser.

Bring two sounds into the light

Start instantly with the two synthetic vowel studies.
Sample A
Warm vowel study Spectral plate revealed
Drop an audio file here, up to 25 MB. The first 20 seconds are analyzed.
Duration 0:02
Sample B
Bright vowel study Spectral plate revealed
Drop an audio file here, up to 25 MB. The first 20 seconds are analyzed.
Duration 0:02
Two spectral plates are ready
Frequency ceiling
Sample A Sample B Estimated formant guides Brighter ink means stronger energy

Read the resonant fingerprints

These values describe average spectral peak positions across voiced frames. Differences are measurements, not a similarity score or identity conclusion.

F1 First resonance region
A
Not available
B
Not available
Gap
Not available
F2 Second resonance region
A
Not available
B
Not available
Gap
Not available
F3 Third resonance region
A
Not available
B
Not available
Gap
Not available

Educational signal visualization only. Formant guides use smoothed spectral peak regions, not a validated LPC workflow, and must not be used for speaker identification.

Zoom 100%
Utilities Studio

Want this utility on your website?

Customize colors and dark mode for WordPress, Notion or your own site.

Frequently Asked Questions

What does a voice spectrogram show?

A spectrogram maps time horizontally, frequency vertically, and signal intensity through color brightness. Sustained speech resonances often appear as horizontal energy bands.

Are my audio recordings uploaded?

No. Audio decoding, spectral analysis, visualization, and playback happen locally in the browser. The tool does not send the selected files to a server.

What are the F1, F2, and F3 guides?

They are educational estimates of three broad spectral envelope peaks. F1 and F2 are commonly used to discuss vowel height and vowel place, while F3 can reflect additional vocal tract resonances.

Can this analyzer identify a speaker?

No. Visual resemblance or formant proximity cannot establish identity. Forensic voice comparison requires validated methods, suitable recordings, uncertainty assessment, and qualified expert interpretation.

Why can formant estimates change with the ceiling?

The selected frequency range changes which spectral peaks are available and how they are separated. Speaker anatomy, vowel, recording quality, pitch, and analysis settings also affect estimates.

# How a Voice Spectrogram Turns Sound Into a Visible Landscape

A voice spectrogram transforms a recording into a map with time on the horizontal axis and frequency on the vertical axis. Stronger energy appears as brighter color. This makes sustained vowels, harmonics, silence, noise, and changing resonances easier to explore than they are in a waveform alone.The analyzer divides the signal into short overlapping frames, applies a Hamming window, and transforms each frame from amplitude over time into energy by frequency. A short frame preserves when a sound happened, while its frequency bins reveal where energy is concentrated. Because every spectrogram balances time resolution against frequency resolution, narrow transients and stable vowels can never both be represented with unlimited precision. The display should therefore be read as a measured view of the chosen settings, not as a perfect picture of the original pressure wave.

Private Browser Processing

Worth noting
The selected recording is decoded into an in-memory audio buffer and analyzed locally. No upload is needed, so private classroom recordings, rehearsals, and personal voice samples remain on the device.
Time Read from left to right
Hz Frequency position
Energy Shown as luminous intensity

# Reading Formants Without Overstating the Result

Formants are resonant regions shaped by the vocal tract. F1 and F2 are often used in phonetics to discuss vowel height and place. This analyzer traces broad smoothed peaks in three frequency regions so beginners can connect visible bands with approximate F1, F2, and F3 behavior.Professional formant measurement normally uses a carefully configured linear predictive coding workflow, checks the tracking by eye, and adapts the formant ceiling to the speaker and vowel. Pitch harmonics, nasalization, room reflections, lossy compression, background noise, and weak recording levels can all pull a simple peak estimate away from the vocal tract resonance of interest. The guides here intentionally expose broad regions and average values for learning. If a guide jumps between bands or conflicts with the visible spectrum, treat that disagreement as a reason to inspect the recording and settings rather than as hidden evidence.
Guide Search region Useful interpretation
F1180 to 1000 HzA broad first resonance region often associated with vowel openness
F2900 to 3000 HzA broad second resonance region often associated with front and back vowel position
F32000 to 4500 HzA higher resonance region affected by vocal tract shape and articulation

# Why Analysis Settings Change the Picture

Lower Ceiling

A 4 kHz ceiling gives more visual space to lower speech frequencies.

  • Useful for a close look at lower resonances
  • May exclude higher energy
  • Does not guarantee more accurate formants

Higher Ceiling

A 6 or 8 kHz ceiling includes more upper spectrum detail.

  • Useful for brighter voices and broadband sounds
  • Shows frication and upper harmonics
  • Compresses lower bands vertically

# A Responsible Two Sample Comparison

Comparing two plates is most useful when the recordings contain the same vowel, word, or short phrase and were made with similar microphones and environments. The displayed gaps are absolute differences between average peak positions. They do not model within speaker variability, between speaker variability, channel mismatch, speaking style, health, age, or the probability of competing explanations. For that reason the analyzer never converts a gap into a match percentage, identity badge, or forensic conclusion.
  • Match the spoken material: repeated vowels or words are easier to compare than unrelated phrases.
  • Use similar recording conditions: microphones, compression, background noise, and distance can alter the spectrum.
  • Listen with the cursor: connect a visible event to the exact moment that produced it.
  • Avoid identity claims: a similar looking spectrogram does not prove that two recordings share a speaker.

What This Analyzer Is For

Generate an audio spectrogram locally from common browser decodable files.
Explore two samples in mirrored or parallel plates with synchronized playback.
Learn how spectral energy and approximate formant regions change across a recording.
Keep comparison descriptive and educational rather than forensic or biometric.