Super Sonique / listening room

Less noise.
More voice.

Super Sonique meets Sidon. Listen to 40 real-noise clips and 10 artificially noised clean clips, side by side with their matching regular Auphonic references.

Non-Studio EDM step 20,000 · EMASampling Heun 8 · 15 evaluationsAudio 48 kHz · full clipsReferences Regular Auphonic · non-Studio
How to read the scores

MOS score recovery.

Recovery = 100 × (denoised MOS − noisy MOS) / (reference MOS − noisy MOS). Scores are predicted by Microsoft DNSMOS P.835, not human ratings. Negative recovery means degradation; above 100% means a score above the reference. A clean–noisy gap below 0.1 is marked n/a. Summary percentages use cohort mean scores, not the mean of clip percentages.

Ground truthMatching regular Auphonic reference
NoisyReal codec baseline / actual synthetic input
Super SoniqueNon-Studio EDM · EMA step 20,000
Sidon v0.1Original Sidon, not DialogueSidon

Loading the listening room…

Real recordings

More real-noise clips

The remaining test pairs, ordered by Super Sonique MOS recovery. The noisy listening baseline is a KVAE reconstruction, as in the previous site; both denoisers receive the original noisy waveform. MOS is calculated on the public MP3s.

Controlled corruption

Artificial noise · +5 dB

Ten held-out regular clean speech clips, mixed with recorded noise using the Sidon augmentation pipeline. Noise comes from the training noise bank: this tests held-out speech under a matched noise distribution, not unseen noise. No reverb, bandwidth, codec, clipping augmentation or packet loss is enabled. The mixer’s final saturation still applies; +5 dB is measured before saturation.

The noisy player is the actual corrupted input. Following the evaluation guide, MOS uses pre-playback float WAVs for this section. Expand codec diagnostics to hear noisy and clean KVAE reconstructions. Best Super Sonique MOS recovery first.