Speech emotion recognition. Record or upload a short clip of speech and a Whisper-tiny encoder, distilled from a model eleven times its size, reads the emotion from how it is said: angry, bored, disgust, fear, happy, neutral or sad. It all runs here in your browser, so your voice never leaves the page.
EMO-DB actors read neutral sentences in each emotion, so the model has to go by how something is said, not by the words. It learned from ten German actors in a quiet studio, so on other voices, languages and microphones it is much less reliable: calm speech tends to come out sad, disgust or bored. Speak one expressive sentence of a few seconds, close to the microphone. You can also drop an audio file here.
Web Audio decodes the clip; a JavaScript resampler, fitted to the soxr filter used in training, brings it to 16 kHz mono. A 400-sample Fourier transform every 10 ms gives the power at each frequency, and 80 mel filters fold it onto a scale that follows how we hear pitch, in decibels down to 80 dB below the loudest cell. That is Whisper's own input; check_mood.py matches it to Hugging Face's feature extractor.
The encoder of OpenAI's Whisper-tiny, pretrained on 680,000 hours of speech, turns the spectrogram into one vector per 20 ms. A learned mix of its layers is pooled by attention (the focus curve under the spectrogram shows the weights) and a linear layer scores seven emotions. It was fine-tuned on EMO-DB while imitating a fine-tuned Whisper-small, a teacher with eleven times the parameters.
Each of the ten speakers was held out in turn, and every choice of model and recipe was made without that speaker: on voices it never heard, the model averages 83.8% recall per emotion, against 77.5% for the previous ResNet-18 and 90.6% for the teacher. On the original split's 54 test clips, picked on validation alone, it gets about 93% right (ResNet-18: 87%). Those figures hold for EMO-DB's own voices and studio: on neutral English audiobook speech it almost never says neutral, and below 70% confidence, where new voices come out right only half the time, the page says “not sure”. It runs in ONNX Runtime Web; the 16 MB model downloads once, with your first own clip.