The problem in one sentence
Music from a generator arrives loud and full range. Voice from a podcast or text to speech tool usually does not, and on a station the two play back to back.
The result is the thing every new station owner hears on their first show: the theme and the song sound great, then the talk segment arrives and it is quiet and muffled, and every listener reaches for the volume knob. Nothing is broken. The files were simply made to different standards.
What the numbers should be
Two measurements matter, and they fail in different ways.
Loudness, measured in LUFS. Aim for -14 LUFS for everything you put on air. That is the streaming standard and it is what music generators target. A talk segment at -25 LUFS is not slightly quiet, it is about eleven decibels down, which is close to a quarter of the perceived volume of the track before it.
Sample rate. Use 44.1 kHz or 48 kHz, stereo. This one is unforgiving: a 16 kHz file has no content above 8 kHz at all, so it sounds like a phone call. Turning it up does not help. The detail was discarded when the file was made and no amount of processing puts it back.
Real numbers from a Studio Reviews episode, which is exactly the trap:
Intro theme 48000 Hz stereo 187 kbps -14.2 LUFS good
Music track 44100 Hz stereo 256 kbps -13.9 LUFS good
Podcast export 16000 Hz mono 48 kbps -25.0 LUFS both problems
The theme and the track sit within half a decibel of each other. The podcast is eleven decibels down and missing the top half of its frequency range.
Fix it at the source first
Check your export settings before you check anything else. Podcast and text to speech tools often default to a low sample rate and mono to keep files small, and many let you change it. Re-exporting at 44.1 kHz stereo takes seconds and is the only way to recover the missing bandwidth.
If the tool cannot export higher, the file will always sound duller than the music around it. You can still fix the volume, which is most of the perceived problem, but plan for the difference.
Fix the loudness
Raising the level is straightforward and safe. Any of these work:
- Any audio editor with a normalize to -14 LUFS option
- A loudness normalization plugin on the export
- The mastering tools on AetherWave, which target broadcast levels
Two rules when you do it:
Do not change the length. Your show's timeline is built from segment durations. If a file comes back even a second longer or shorter, every segment after it shifts, and the show will not end where you think it does.
Leave headroom. Target a true peak around -1.5 dB rather than pushing to 0. A file that clips will distort on some devices and sound fine on others, which is the worst kind of problem to chase.
Check it before you schedule it
The fastest check costs nothing: play your talk segment, then immediately play one of your music tracks, without touching the volume. If you flinch or reach for the knob, so will every listener. That comparison catches both problems at once and needs no tools.
For a real measurement, any loudness meter will report LUFS. Aim for -14, accept -16 to -13, and treat anything below -20 as needing work.
Why this matters more on a station than anywhere else
A podcast played on its own sets its own level. The listener adjusts once and forgets.
A station does not work that way. Your segment is sandwiched between mastered music, on a stream the listener is not watching, often in a background tab. They will not investigate a quiet segment. They will assume the station is broken, or turn it off.
The station is only as consistent as its least consistent file.