How to mix vocals, and why your mix sounds muddy

Four problems diagnosed in a fixed order. Why muddiness has a specific frequency address, the fix that is subtraction rather than addition, and what changes when the vocal came from a generator.


A vocal that sits badly in a mix almost always has one of four problems, and they get diagnosed in a fixed order. Fix them out of order and you spend an hour on reverb when the real issue was that the bass and the vocal are fighting for the same frequencies.

Here is that order, plus what changes when the vocal came out of a generator rather than a microphone.

The order of operations

1. Level and dynamics. Before any tone shaping, the vocal needs to sit at a consistent volume. A performance that swings twelve decibels between quiet and loud phrases cannot be made to sit in a mix by any amount of EQ. Compression, or manual level rides, first.

2. Subtractive EQ. Remove what does not belong before adding anything. Low rumble below roughly 80 Hz on most voices, and any specific resonance that jumps out. Cut narrow, boost wide, is the old rule and it holds.

3. Space for the vocal in everything else. This is the step people skip. A vocal sounds buried not because it is too quiet but because something else occupies its range. Usually guitars, synth pads or a busy midrange. Duck those rather than pushing the vocal louder.

4. Only then, effects. Reverb, delay, saturation. These place the voice in a space, and placing something you have not yet balanced is decorating a room with no floor.

Doing these in reverse is the single most common amateur mixing mistake.

Why your mix sounds muddy

Muddiness has a specific address: too much energy between roughly 200 and 500 Hz, contributed by several sources at once.

The vocal has body there. So does the guitar. So does the piano's left hand, the snare's fundamental, and the room tone of anything recorded live. Individually each is fine. Together they sum into a blanket.

The fix is subtraction, not addition:

Cut, do not boost. If the vocal is unclear, cutting 250 Hz on the guitar usually beats boosting 3 kHz on the vocal. The second approach makes the vocal harsh; the first makes room.

High-pass generously. Almost nothing except bass, kick and the lowest body of a vocal needs content below 100 Hz. Filtering it out of everything else clears space you did not know you had.

Solo less. Everything sounds better soloed and mixes are heard in context. Judge the vocal against the full arrangement, always.

Check on a phone speaker. A small speaker reproduces almost nothing below 200 Hz, which is exactly why it exposes an arrangement that only works on headphones.

The four vocal problems, and their fixes

Buried. Something else is in its range. Duck the competitor, do not raise the vocal.

Harsh. Usually 2 to 5 kHz. A gentle wide cut, or a de-esser for sibilance specifically, which lives higher around 5 to 8 kHz.

Thin. Missing body around 150 to 250 Hz, often because it was high-passed too aggressively. Add narrowly and listen for mud returning.

Disconnected from the track. The voice sounds pasted on top rather than in the room. This is a space problem: a short reverb or a slap delay that matches the arrangement's ambience glues it. Saturation helps too, because a little harmonic distortion makes a voice feel recorded rather than inserted.

Mixing a generated vocal, which is different

Three differences worth knowing.

You may not have a separate vocal track. A generated song usually arrives as a stereo mix, so the only route to vocal-specific processing is separation first, with the artefacts that brings. Stem separation is the enabling step and it is not free.

The vocal is already processed. Generated audio arrives with compression, reverb and tone shaping baked in. Stacking your own on top is where things get brittle. Start with far less than feels right.

Most vocal problems are prompt problems. A thin, harsh or oddly placed voice is usually easier to fix upstream than downstream. Tokens like characterful male lead, breathy female vocal, close-mic intimacy, raw live tracking and wide dimensional stereo shape the voice at generation time, and every recipe in our catalogue carries a vocal decision explicitly. Controlling the voice in a prompt covers what each token does, and general output quality covers the rest.

The practical rule: try one more generation with a sharper vocal token before you reach for a de-esser on a separated stem.

A checklist that fixes most vocals in ten minutes

  1. Level the performance so the loudest and quietest phrases are within a few decibels.
  2. High-pass the vocal, and everything else that does not need low end.
  3. Find and cut the one resonance that jumps out. There is usually exactly one.
  4. Cut 200 to 400 Hz on whatever else is crowding the vocal.
  5. Add a short reverb, less than you think, matched to the arrangement.
  6. Check on a phone speaker.
  7. Leave it, come back tomorrow, and listen once before changing anything.

That last step catches more bad decisions than any plugin.

FAQ

Why does my vocal sound buried? Something else occupies its frequency range. Duck the competing instrument rather than raising the vocal.

Why does my mix sound muddy? Too much accumulated energy between 200 and 500 Hz from several sources. Cut it on the supporting parts rather than boosting the vocal.

Should I compress before or after EQ? Level and dynamics first, then subtractive EQ, then space for the vocal in the arrangement, then effects.

How much reverb on a vocal? Less than sounds right when soloed. Judge it in the full mix only.

Can I mix a vocal from an AI-generated song? Only after separating stems, which introduces artefacts. Fixing it in the prompt is usually cheaper and cleaner.

What is the fastest improvement I can make? High-pass everything that does not need low frequencies. It takes two minutes and clears the most space.

Keep reading