Harsh S sounds on AI vocals: de-essing without lisping
Updated 2026-09-06
An S that cuts through a track like a needle is one of the most common complaints about generated vocals, and one of the easiest things to overcorrect.
Where the S lives
Sibilance is a burst of energy between roughly 5.5 and 9.5 kHz. Generators tend to produce it a few dB hotter than a recorded vocal because the model learned bright, close-miked pop vocals and reproduces the brightness without the compression that tamed it on the record.
How a de-esser works
It is a compressor that only listens to that band. When the S arrives, it turns that band down for a few milliseconds and lets go. The rest of the vocal is untouched. A static EQ cut at 7 kHz would do the same job on the S and also dull every word between the S sounds, which is why de-essing exists.
How much to take
- 2 to 4 dB of reduction on the loudest S sounds is normal.
- Past 6 dB the singer starts to lisp. "Sun" becomes "thun." That is the sign to back off.
- Do it before the limiter. A limiter reacts to sibilance first and pumps the whole mix on every S.
- Listen on earbuds. Sibilance shows up there first.
What Burnish does
Every genre profile carries a de-ess amount tuned for its typical vocal, and the de-ess control scales it from off to full. The report prints the band it worked on and how much it took. If a track has no vocal, the de-esser finds nothing to do and the report says so.
Check your own file, free
Drop a track in and get the cutoff, the low-mid excess, the width, the loudness and the true peak on one page. Nothing is changed and nothing is charged.
Run the free analysis Get Burnish for iPhone