How do you get the words out of a song you only have as audio?

Three different problems hide behind that question, and automatic recognition solves only one of them. The other two need a search box or a blank page.


Start with the file, not with the tool

A voice memo. Three minutes and eleven seconds, recorded in a parked car with the engine off. The guitar is too close to the mic, the voice sits slightly behind the beat, and somewhere in the second half there is a line you were pleased with at the time and can no longer make out.

You want that line as text. Fair enough. But which tool helps depends on a question almost nobody asks first: whose words are these, and do they exist yet? Three completely different problems arrive wearing the same sentence, "I need the lyrics out of this audio", and only one of them is solved by pointing software at the file.

The three cases people keep mixing up

Case one: you recorded it, and you want the words on paper

You sang them, they are on the file, and now you want them in a document so you can rework the second verse or print a sheet for whoever is singing. This is a transcription job, and it is the only one of the three where automatic recognition is the right first move.

Case two: it is somebody else's song, and you want to sing along

On the surface this looks like case one. It is a lookup problem instead. Someone has almost certainly typed those words out already, checked them against the master, and argued in the comments about the one word in the bridge. Running a recogniser over a commercial mix to rediscover that alone is the slow route to an answer that exists as text.

One sober line before moving on: writing out a song you love so you can sing it correctly in the kitchen is your own business, and putting another writer's words out under your name is a different act entirely.

Case three: there are no words yet

The file is a melody with mumbling on top. na na na, something something, and the thing about the door. Feed that to a recogniser and it will return a transcript, because that is what it does. The transcript will be nonsense, and the nonsense is accurate: it is faithfully reporting that you did not sing any words. Nothing is broken. There is no text in the file to extract.

This is where most people burn the most time, because the tool produced an output, and an output feels like progress.

Where we fit, said straight, before you read further

sunomarket does not listen to audio. We do not transcribe, we do not split stems, and we do not produce sound of any kind: no vocal, no instrumental, no mp3 at the end. What our lyrics generator does is write words. You give it a theme, pick from a handful of lists, and it returns a title and a finished lyric for 20 credits, which is $0.20 at the rate of 100 credits per dollar. If you want those words sung, you take them to Suno or another service that makes audio. Suno is a separate company, we are not affiliated with it, we do not resell it and we do not control its rules.

For case three we are the answer. For cases one and two we are not the tool, and the next few sections are about the tools that are.

Why machines hear speech well and singing badly

Speech recognition got good on a diet of podcasts, phone calls and meeting recordings. Singing breaks most of the assumptions that made it good, and it breaks them all at once.

The music lives in the same frequency range as the voice. A human voice does its intelligibility work roughly where a snare, an acoustic guitar and a piano's middle register also live. Your ear separates them because it knows what a voice is. A recogniser working on a stereo mix gets one stream with all of it stacked together, and the t at the end of a word arrives at the same instant as a hi-hat.

Vowels get held past the point where they carry information. Spoken, "stay" takes a fraction of a second. Sung, it can run for several bars, and most of that is a held vowel with no consonant to say which word it belongs to. Recognisers built on speech timing read that stretch as extra syllables, so a four-word line comes back as seven words.

Consonants get softened on purpose. Good singers round off hard endings because a spat t sounds ugly over a sustained chord. Every dropped consonant is evidence removed from the exact place a recogniser looks first.

Breath is part of the performance. Breathy delivery, the thing that makes an intimate verse feel intimate, is noise laid directly over the signal. The quieter and closer the vocal, the more of the word is air.

There is often more than one of you. A doubled lead, a harmony stack, a chorus sung three times and layered. Two versions of the same phrase a few milliseconds apart do not average into a clearer phrase. They smear into one.

Reverb and delay glue the words together. The tail of one word is still sounding when the next word starts. That is usually a production choice you made deliberately, and it is exactly the boundary a recogniser needs to find.

A good number of the words are not words. Names, invented spellings, ooh and mm used as rhythm rather than meaning, a street in your town, slang that entered the language last year. Anything outside the vocabulary gets replaced by the nearest thing inside it, confidently.

The model expects grammar, and lyrics do not supply it. Recognisers lean on a language model to choose between similar-sounding candidates, and that language model was trained on sentences that behave. Lyrics invert word order, drop subjects, repeat a fragment four times and end on a preposition. The machinery that makes speech transcription accurate pulls singing toward plausible prose and away from what you actually sang.

These stack. A breathy doubled vocal with reverb, singing a made-up name over a busy mix, fails on five of them at once. That is not a broken tool. That is a tool being asked to do a different job than the one it was built for.

Raising the odds, when the recording is yours

This part applies to case one only, because it needs access to the material. Ordered by how much each step actually buys you.

1. Use the vocal track, not the mix. If the session still exists, bounce the lead vocal alone, dry, with reverb and delay bypassed and heavy compression off. This one step does more than the other five combined, because it removes the two biggest obstacles in a single move. Solo the take you actually used and mute the doubles.

2. No session? Split the stems. Vocal isolation tools pull the voice out of a finished mix as a separate file. The result carries artefacts, a watery ringing around the edges of words, and it is still far easier to recognise than the full mix. Know the catch going in: isolation works best when the vocal is loud and centred, and worst on exactly the material you most want help with, a quiet vocal buried in a dense arrangement.

3. Cut the file into sections. Recognisers drift: a four minute file transcribed in one pass tends to be sharp early and increasingly inventive later. Feed it a verse at a time as separate short files, and you also get a transcript already split by section, which you wanted anyway.

4. Slow it down. Half speed, without pitch shifting if the tool offers that. Fast syllabic material, a rap verse or a patter bridge, benefits far more than a slow ballad does. On a ballad it can make things worse, because holding an already long vowel twice as long gives the recogniser even more empty runway to invent syllables in.

5. Handle doubles and harmonies as their own pass. If the chorus is three stacked takes, transcribe one of them by itself. The words are the same. You need one clean reading, not an average of three.

6. Run it twice and compare. Two passes, or two different tools, over the same section. Where they agree, the odds are good. Where they disagree, you have a shortlist of the lines that need your ears instead of a whole song.

These have sharply diminishing returns. Steps one and two change the outcome. Step six mostly tells you where to look.

Fix the result by rhyme and meter, not by meaning

Here is the trap. A recogniser's mistakes are built to be plausible. It does not output gibberish, it outputs a grammatical sentence made of words that sound similar, which means reading the transcript for sense will not catch the errors. Sense is precisely what the machine optimised for.

Rhyme and syllable count are different. They are countable, and the machine was not trying to preserve them.

Say verse one, line four, ends like this:

nobody taught me how to wait

Eight syllables: no-bo-dy-taught-me-how-to-wait. It lands on the sound of wait.

The transcript hands you this for verse two, line four:

nobody talks about the weight of it

Read it for meaning and it passes. It is a sentence, it fits the subject, you would nod at it. Count it and it falls apart: ten syllables, and it ends on it, which chimes with nothing. Verses in one song generally hold the same shape, and this one has grown two syllables and lost its landing.

The real line was:

nobody talks about the weight

Eight syllables, ending on the same sound as wait in verse one. The recogniser heard you hold the vowel in weight, read the held vowel plus your breath as two more syllables, and filled them with the most probable English words that could follow. That failure is systematic rather than random, and once you know it you will find it in every transcript you run.

So work like this. Put the transcript next to the audio and check three things per line, in this order: the syllable count against the matching line in the other verse, the last stressed vowel against whatever it is supposed to rhyme with, and only then the meaning. If you are hazy on which line is answering which, the guide to rhyme schemes covers it. And if you are transcribing a song in order to learn from how it was built, the walkthrough on analysing a song is a better frame than a bare transcript, because it tells you what to look at once you have the words.

Two habits worth keeping. Bracket every uncertain word instead of guessing, because a confident wrong word will still be sitting there in six months pretending to be right. And read the transcript aloud once, in time with the track, which catches what silent reading never does.

When none of this works

The advice above has a ceiling, and it is worth naming.

Screamed, growled or heavily distorted vocals. The information a recogniser uses is mostly gone before the signal reaches it. Expect very little, whatever you do to the file.

Heavy pitch correction, vocoder, formant shifting, stacked octaves. Processing that turns a voice into an instrument turns it into an instrument for the recogniser too.

Two languages inside one line. Recognisers pick a language and commit, so a line that switches halfway comes back wrong on one half. Transcribe the languages as separate passes where the sections are separate. If the switch happens mid-line, that line is a manual job. Writing songs in another language deals with the composition side of this.

A phone recording with wind, a live room, or a table tap on the mic. Nothing downstream fixes a signal that was never captured.

Proper nouns, invented words, and a chorus that is deliberately a sound rather than a sentence. Those come back wrong every time. Type them in yourself and move on.

In all of these, accept a partial transcript. Take the share the machine can give you, then do the hard lines by ear at half speed. Ten hard lines is a twenty minute job, and fighting a tool for an hour to avoid those twenty minutes is a bad trade.

Case three in detail: nothing to transcribe yet

Back to the mumble demo, because it is the one with a genuinely different answer.

What you have is more valuable than it looks. A hummed take already carries the melody, the phrasing, the number of syllables per line, where the line breaks fall and where the hook lands. That is a lyric-shaped hole with exact dimensions, and it contains no words at all, so there is nothing to recognise and no transcription tool on earth will help. What helps is writing the words, either yourself or by reworking a draft.

Here is the honest version of what our tool does in that situation. The lyrics generator takes one free text field for your theme, capped at 300 characters, which is about three sentences. Everything else is a menu: point of view, five options. Song language, nine. Audience age, sixteen bands. Structure, six templates. Which direction the hook pushes, five. Vocal, three options, and there is no duet among them, so two voices trading lines is something you arrange yourself afterwards. Musical style, 53 curated root styles. Emotion, one to three picks from twelve. It returns a title and a lyric, up to 20,000 characters, for 20 credits, that is $0.20. Current rates live on the pricing page.

What it does not take is your audio. It cannot hear your melody, it does not know your second line is seven syllables, and it will not match your phrasing by itself. The workflow therefore has a manual step in it, and pretending otherwise would be a lie.

A workable order of operations for a hummed demo:

  1. Before generating anything, write down what your demo already decided. Syllables per line for each section, where the rhymes land, which line is the hook. You get this by counting your own mumbles. It takes five minutes and it is the part nobody does.
  2. Spend the 300 characters on the specific thing rather than the category. my brother moved out at 19 and calls on Sundays, we talk about nothing, I want to say the thing I never say beats a song about family by a wide margin, and both fit inside the limit.
  3. Pick the structure template that matches the shape your demo already has, not the one you like on paper. The melody is already committed.
  4. Generate. At $0.20 a run, a second and third run with a different emotion set or a different hook direction costs forty cents, and comparing three drafts teaches you what your song wants faster than staring at one.
  5. Now do the fitting by hand. Sing the generated lines against your demo and cut or add syllables until they sit. Your melody wins every argument. A line that reads beautifully and does not fit the tune is the wrong line.
  6. Once the words are settled, carrying them into a music service is the separate step covered in turning lyrics into a song.

Step five is where the work is. A generated lyric hands you images and phrasings that are on-theme, and the shaping against a real melody stays yours. The same is true of a lyric a co-writer hands you, so this is not a complaint about generation, it is how the job goes.

Questions people actually ask

Can I get lyrics out of a song using your generator? No. It does not accept audio and it does not listen. It writes new words from a theme and a set of choices, for $0.20. If you have audio with words already in it, you need a transcription tool, and the middle sections of this article are about getting the most from one.

Is there a tool that does this perfectly? Not for singing over music. A single dry vocal, clearly enunciated, in a widely spoken language, transcribes well; a doubled breathy vocal buried in a mix does not. Judge any tool on your hardest file, not your easiest.

The transcript is grammatically perfect and completely wrong. Why? Because that is the failure mode, not an accident. A recogniser resolves ambiguity by picking what is most likely as language, so its errors come out fluent. Check by syllable count and rhyme, which it was never optimising for.

My demo has no words. Which field do I put the melody in? There is no field for it. The generator writes to a theme and your menu picks, and matching the result to your melody is a manual pass afterwards. Write your syllable counts down first so the fitting takes minutes instead of an evening. If you want words that come out singable from the start, the guide to writing song lyrics is about that craft directly.

Two singers trading lines, can the generator do that? Not as a setting. Three vocal options, no duet among them. Generate a single-voice lyric and split the lines between the two parts yourself.

What if my song is in a language the recogniser handles badly? Then it will guess in that language's nearest common neighbour and the result will be unusable. Transcribe by ear. And if you are writing rather than transcribing, set the song language in the generator directly instead of writing in one language and translating afterwards, because translated lyrics almost always lose their meter.

Keep reading