Noise removal for podcast recordings
A guest on a laptop mic in a room with a fridge. An interview recorded in a café because that is where they agreed to meet. This page is about where noise removal belongs in an episode workflow, what to feed it, and what it will not fix.
Clean each track, not the mix
If your recorder or remote platform gave you a separate file per microphone, those files are the input — not the bounce. Three reasons, and none of them is that the model is smarter about single tracks:
- One preset per run. Your co-host in a treated room and your guest on a kitchen laptop do not need the same amount of processing. In a mix you can only pick a compromise; per track you can leave one alone and push the other.
- The output is mono. Everything is summed to one channel before the model runs, because the model is single-channel. Feed it a stereo mix and the stereo goes away. Clean the tracks, then mix — the image survives.
- It is not reversible. Bake noise removal into the bounce and every later decision is made on top of it. Per track, you still have the original next to the cleaned version.
Do this before you set levels
We do no loudness normalisation, and taking noise out takes energy out: on our own sample the cleaned version came back about 2 LU quieter than the original. So compression, levelling and whatever you use to hit your delivery target all belong after this step. Level a noisy track first and you have made every decision against a noise floor that is about to move.
Music and effects belong after it too. The model keeps what it recognises as speech and pulls down the rest, so a bed under the voice is precisely the sort of thing it is built to remove. Clean the speech, then lay the music in on top.
What it will not fix
- Another conversation in the room. This is a speech model — it keeps speech, including speech you did not want. The café table behind your guest is close to the worst case for it, not the easy one.
- Room echo. A guest in an empty bedroom is hearing their own voice arrive late, which is a different problem from noise added on top. We have not measured how much of it this removes; assume it stays and check on the player.
- Clipping and distortion. If the mic was slammed or the gain was pinned, the recording is missing information. Noise removal takes things away; it cannot put back what was never captured. More on where the limits come from.
Testing it on an episode
The preview picks the loudest 30 seconds of whatever you drop in, which is a reasonable guess at a representative passage but is not the passage you are worried about. If there is a specific two minutes where the air conditioning kicked in, export just that and test on it — you will learn more from your worst stretch than from your best one.
Start on Balanced. If you go to Strong, listen for breaths, sibilants and the tails of words, because that is where over-processing shows up first as a watery or robotic edge. Clear Voice does less and leaves more of the room in, which is often the right answer when the noise is mild and the voice matters more than the silence. Switching presets re-runs in seconds, so try all three before deciding.
Be clear about what today's version is for, though: it makes a 30 second preview, and cleaning a complete file is not built yet. For a full episode this is a test of whether the approach works on your recordings — not yet a way to deliver one.