Why does Suno scream on every track, even when I ask for a quiet song?
The complaint everyone has
"Every track has screaming vocals no matter what I put." "Pitch escalation, melisma and screaming regardless of what you prompt." "It just decides to escalate the vocal intensity toward the end."
This is real, and it is the most reported vocal problem with Suno. You write quiet vocals, calm, soft, and by the last chorus the singer is belting like the building is on fire. Here is why, and here is the fix.
Why quiet fails
The model treats a lot of your prompt as genre energy, not as a per-second volume rule. Words like emotional, powerful, epic, anthem, chorus, even build push the model toward escalation. A single quiet sitting next to them loses the fight. The model also has a strong prior that songs grow — verses are calm, choruses are big — so left alone it will crank intensity toward the end no matter what one adjective says.
There's a deeper reason too. Mood words like quiet, soft, calm don't map onto dynamics at all — Suno effectively ignores them, because they describe how the listener should feel, not what the voice should do. The model can't act on a feeling. It can act on a technique. So don't argue with it using one weak word. Direct the performance, and pin the dynamics section by section.
Describe the manner, not the mood
Instead of a mood word (calm), describe the physical vocal technique. Technique tags are things the model can actually execute:
[Vocal: restrained, breathy, close-mic, conversational]
[soft head voice]
no belting
spoken-word delivery, almost whispered
low dynamic range, steady
These work because they describe how the voice is produced, not how the listener should feel. "Breathy" and "close-mic" are incompatible with screaming — you can't scream on a close mic without clipping, and the model knows that. On v5.5 especially, a detailed delivery anchor holds better than a vague one: the more concrete the technique, the harder it is for the model to drift off it.
The translation table: mood word → delivery tag
Every mood word you're tempted to write has a concrete delivery equivalent. Use the right column — it names a technique the model can perform, not a feeling it will shrug at.
| You wrote (mood) | Write this instead (delivery / technique) |
|---|---|
| calm | restrained, breathy, close-mic, steady |
| quiet | low dynamic range, soft head voice, near-whisper |
| soft | breathy, unpushed, conversational |
| gentle | soft head voice, minimal vibrato, legato |
| intimate | close-mic, whispered verses, dry vocal |
| tender | understated, warm, restrained delivery |
| powerful | full-chest tone, forward in the mix, sustained (still no belting if you want control) |
| raw | dry vocal, no reverb, single close mic |
| mellow | relaxed phrasing, laid-back timing, low intensity |
Notice what the right column never contains: an adjective about emotion. Every entry is a mixing or vocal-coaching instruction. restrained and close-mic box out screaming the way a slow tempo boxes out a dance beat — by crowding it out, not by forbidding it.
Pin dynamics per section — the Lyrics field does the real work
Here's the mechanic most people miss: arrangement and delivery written into the Lyrics field are roughly ten times stronger than the same words in the Style box. The Style field sets a global average; the Lyrics field controls the track second by second. Dynamics are a per-section thing, so they belong in Lyrics, tagged section by section.
Don't set the vibe once and hope. Label the intensity inside every structure tag so the model has no room to escalate:
[Verse 1 - soft, intimate, breathy]
...
[Chorus - still restrained, no belting, warm]
...
[Verse 2 - same low intensity, conversational]
...
[Outro - fades to near silence, whispered]
The key move: your chorus is where escalation lives, so that is exactly where you write still restrained, no belting. Repeat the constraint everywhere the model wants to grow.
You can stack constraints inside one tag using the pipe |, which combines metatags cleanly:
[Chorus - restrained | no belting | no melisma | warm]
One section header, four locked constraints, no comma pile for the model to skim. This is the same discipline covered in directorial dynamics — write the arc into the sections, don't name a mood and hope.
The escalation-trigger word list
Go through your prompt and remove the words that provoke screaming. These are the fuel; deleting them does more than any amount of quiet:
- Intensity words:
powerful,epic,explosive,intense,huge,massive - Belt cues:
belting,soaring,wailing,screaming,full-throated,riffs and runs - Arc words:
build,climax,crescendo,drop,bigger chorus,explosive finale - Anthem cues:
anthemic,anthem,stadium,arena,festival - Sneaky ones:
emotional(swap fortender/understated),passionate(swap forintimate),energy(swap forsteady)
You are not adding "quiet." You are removing the fuel. A prompt with none of these words gives the model nothing to escalate toward.
Two working calm-vocal recipes
Recipe 1 — intimate acoustic ballad
Style box:
slow acoustic ballad, 70 BPM, close-mic vocals, restrained, breathy, conversational, warm, low dynamic range, no belting, intimate
Lyrics header:
[Intro - soft fingerpicked guitar]
[Verse 1 - whispered, close, breathy]
[Chorus - warm, restrained, no belting, no melisma]
[Verse 2 - same intimacy, spoken-word feel]
[Outro - fades to near silence]
Recipe 2 — spoken-word lo-fi (maximum control)
When even a restrained sung chorus creeps up in intensity, drop the singing altogether and anchor on spoken delivery — it has no ceiling to climb toward.
Style box:
lo-fi bedroom pop, 82 BPM, dry close-mic vocal, spoken-word, unprocessed, low dynamic range, tape hiss, no drums swell
Lyrics header:
[Verse 1 - spoken, close-mic | breathy | flat delivery]
[Chorus - half-sung | still spoken-feel | no belting | no melisma]
[Verse 2 - same restraint | conversational]
[Outro - trails off | whispered | near silence]
Recipe 3 — held-back indie folk
Style box:
indie folk, 75 BPM, warm female alto, understated, close-mic, soft head voice, steady, no belting, sparse arrangement
Lyrics header:
[Verse 1 - soft head voice | minimal vibrato | intimate]
[Chorus - stays understated | no soaring | warm | held notes]
[Bridge - near-whisper | dry vocal]
[Outro - fades, breathy]
Generate two or three times for any of these. Suno is probabilistic, so keep the seed/style that behaves and reuse it — change only one field at a time when you tune, and see the guide on reproducible results.
Do this now
- Delete every escalation word from your prompt using the trigger list above.
- Replace mood tags with delivery tags from the translation table:
breathy,close-mic,no belting. - Write the intensity into every section header in the Lyrics field, especially the chorus — that's where it's ten times stronger. Combine constraints with
|. - Build and test the phrasing in the prompt builder, and browse restrained styles in the catalog.
For the deeper mechanics of shaping loudness across a song, read directorial dynamics. To stop the model changing who is singing between takes, read lock the vocal persona.