2020 portfolio project · Live demo

Good Mood.

Speech emotion recognition. Record or upload a short clip of speech and a Whisper-tiny encoder, distilled from a model eleven times its size, reads the emotion from how it is said: angry, bored, disgust, fear, happy, neutral or sad. It all runs here in your browser, so your voice never leaves the page.

How EMO-DB names a clip
W
Wut
angry
L
Langeweile
bored
E
Ekel
disgust
A
Angst
fear
F
Freude
happy
N
Neutral
neutral
T
Trauer
sad
New voices
83.8%
UAR, 10 unseen speakers
Test clips
92.6%
of 54; ResNet-18: 87.0%
Training data
EMO-DB
535 clips, 10 actors
Encoder
Whisper
tiny, 8.2 M params
Model
16 MB
fp16 ONNX
Runs in
Browser
WASM, no server
01 /

Read the emotion

Speech clip
EMO-DB samples held out from training
EMO-DB samples from the training set

EMO-DB actors read neutral sentences in each emotion, so the model has to go by how something is said, not by the words. It learned from ten German actors in a quiet studio, so on other voices, languages and microphones it is much less reliable: calm speech tends to come out sad, disgust or bored. Speak one expressive sentence of a few seconds, close to the microphone. You can also drop an audio file here.

Predicted emotion
…

    Model input · log-mel spectrogram
    02 /

    How it works

    1 · Spectrogram

    Turn sound into a map

    Web Audio decodes the clip; a JavaScript resampler, fitted to the soxr filter used in training, brings it to 16 kHz mono. A 400-sample Fourier transform every 10 ms gives the power at each frequency, and 80 mel filters fold it onto a scale that follows how we hear pitch, in decibels down to 80 dB below the loudest cell. That is Whisper's own input; check_mood.py matches it to Hugging Face's feature extractor.

    2 · Encoder

    Listen with Whisper

    The encoder of OpenAI's Whisper-tiny, pretrained on 680,000 hours of speech, turns the spectrogram into one vector per 20 ms. A learned mix of its layers is pooled by attention (the focus curve under the spectrogram shows the weights) and a linear layer scores seven emotions. It was fine-tuned on EMO-DB while imitating a fine-tuned Whisper-small, a teacher with eleven times the parameters.

    3 · Measured

    Scored on new voices

    Each of the ten speakers was held out in turn, and every choice of model and recipe was made without that speaker: on voices it never heard, the model averages 83.8% recall per emotion, against 77.5% for the previous ResNet-18 and 90.6% for the teacher. On the original split's 54 test clips, picked on validation alone, it gets about 93% right (ResNet-18: 87%). Those figures hold for EMO-DB's own voices and studio: on neutral English audiobook speech it almost never says neutral, and below 70% confidence, where new voices come out right only half the time, the page says “not sure”. It runs in ONNX Runtime Web; the 16 MB model downloads once, with your first own clip.