How to Work Out Why Your Vocal Remover Result Sounds Wrong

You ran a song through, downloaded the file, and it is not usable. The singer is still hovering in there somewhere. Or the drums went strange. Or the whole thing holds together fine until the chorus arrives and then it does not.

So you try again with a different copy of the song, get a slightly different flavor of wrong, and start wondering whether the tool is any good. Usually the tool is fine. What is missing is a way to read the failure. A bad result is not just noise — it points at a specific cause, and once you can hear which kind of wrong you have, the fix is normally one change rather than five. Everything below assumes a vocal remover working entirely inside a browser tab: you send it a file, it sends one back, nothing lands on your machine.

How to Work Out Why Your Vocal Remover Result Sounds Wrong

Play both files before you judge either one

Almost nobody does this, which is exactly where the guesswork starts.

One pass through a vocal remover produces two files, not one. The instrumental is the song with the lead voice taken out. The acapella is that voice on its own, with the backing gone. Both fall out of the same set of decisions, made slice by slice through the track. Whatever ended up in one of them was, by definition, removed from the other.

They are useful as a couple; that’s why. Play instrumental only and describe as much as you can as “something sounds off. Play both, and you can state where the material is. If the acapella is dragging the cymbals and bass, then this is why the instrumental is sounding hollow — the cymbals and bass are in the other file. This is a mirroring that is properly understood, since both of the other checks below rely on this: The acapella and the instrumental are two sides of the same decision.

Download both. Listen back to back. Then match what you hear against the four patterns that follow.

If the singer is still faintly there

Vowels that remain; words you can almost hear in the chorus, a breath before the line. The model was not totally dedicated.

It’s almost always not about the tool; more often than not, it’s about the original mix. If there’s a guy in the same range as the rhythm guitar or a synth pad, it’s really difficult to separate it because there’s no nice, clean cut in the middle to divide it. Layered harmony makes things worse: a vocal remover is designed to discover one lead voice, and you gave it five voices all singing together.

Look at the acapella first; don’t blame the acapella first. Its tone will be full and confident, if it does. After the song started out patchy, the acapella did as well.

If the backing sounds thin and hollow

When there is a cause, there is an opposite cause. In this instance, it was too far and went too far; real instruments appeared as well as the voice.

The typical cause is the original vocal having a lot of reverb. Reverb tails extend a voice in space (stereo field) and in time; when a vocal remover is tracking, it’s picking up anything else that was in that space — typically the cymbals and the tops of the guitars. The end of your instrumental is not as loud or big as the song that you began with.

The tell is once more in acapella. Fill it with and hear if it has anything loud and clearly not the singer. The easiest indication of a separation that went too far in one direction is if there is percussion bleeding into a vocal track.

If it only falls apart in one section

Verses clean, chorus a mess. Common and can easily be mistaken for a quality issue.

Choruses are denser. More instruments, more vocals, more all competing for the same space. A song that a vocal remover is able to perform without any trouble for two minutes can break down in the 20 seconds that everyone is singing at once.

This builds a good habit: don’t judge the result from the start. Most songs are virtually empty in the first 15 seconds; they split up nicely, and they don’t give you any information about what the actual song is going to be about.

If it sounds fine quietly and bad loud

You tested the laptop speakers; it was all right, then you placed any headphones on and heard the artifacts.

The separation artifacts are generally at the bottom of the spectrum in the low frequencies as small watery smears that are heard up in the high frequencies and as a slight flutter under sustained notes. They are entirely concealed when playing quiet music. So does a noisy room.

My rule here is unarguable: Decide on the rating of the file at the volume it will actually get used at, on the device it will be used at. To rehearse in a rehearsal room and to sleep on a bed under a talking head video have completely different tolerances, and a file that is rejected by one will pass the other.

What to change before you run it again

Three things, roughly in order of how much they matter.

Begin at the beginning. This is the one you’re looking for. If you are feeding a vocal remover an MP3 that you’ve downloaded somewhere, all of the high-frequency detail that the model needs to distinguish voice from instrument is gone prior to your uploading anything. An uncompressed WAV version of the same song will often be able to solve a problem that no re-running will be able to solve. The uploader supports WAV, MP3, M4A or OGG up to 50 MB, if you have a really long uncompressed file, cut it to only the part you want and send that.

Then cut back on your expectations for the job. This isn’t a video bed, so an instrumental doesn’t have to be perfect, as the dialogue will cover up a majority of the things that might bother you. One is required if there is no other music playing, such as during solo rehearsal.

Finally, test properly. The entire process is password-less and allows you to listen to the results before committing to it—there’s no point in making a first impression. Run the same track through two or three times and batch if it takes less time – keep the track that made it through the loud part. Signing in only enters at the time you want the download.

When to stop and pick a different song

There is some material that just isn’t going to work, and you’ll save an afternoon if you know which it is.

With live recordings, the crowd noise that always comes with the live environment usually makes it a lost cause: the voice is baked into the room, and the room is baked into everything else. When voices are the arrangement, it is not possible to extract a ‘lead’ — this is the same as choral and heavy harmony arrangements. If you’re looking for a vocal remover to control drums and bass as separate components completely, know that the vocal remover outputs two halves of the song, not the original session, and that’s a different tool with a different output.

I don’t know, just one more thing which isn’t related to audio quality. The removal of a voice from a song does not imply any rights to the song. A rehearsal file on your own laptop is one; your own video monetized, or a public remix is another; the licensing doesn’t matter how the file was created. Clear that up first before the file departs.

Popular on OTW Right Now!

Add a Comment

Your email address will not be published. Required fields are marked *