How to make a karaoke version of a song with AI
Remove the lead vocal with modern AI separation, keep a clean instrumental, and fix the common problems.
Last updated · 6 min read
A karaoke version is the song without its lead vocal, so you can sing over the instrumental. Ten years ago that meant “center-channel cancellation”, a trick that cut the vocals but also destroyed the bass and drums. Today AI stem separation can pull the vocals out cleanly. Here’s how to do it with MixSongs.
Only use songs you have the right to use. A karaoke version you make for singing at home is a very different thing from one you upload or perform publicly. Check with the rights holders before publishing anything. See our Terms and Copyright policy.
Why AI beats the old “vocal remover” trick
Old vocal removers subtracted the left channel from the right. Vocals are usually mixed in the center, so they cancelled out, but so did the kick, the snare and the bass, and anything with stereo reverb left a ghostly echo. AI models are different: they have learned what a voice sounds like and separate it by its sound, not by its position. MixSongs uses Mel-Band RoFormer, one of the strongest publicly available vocal models, followed by Demucs for drums, bass and the rest.
Step 1: Add the song to MixSongs
Open app.mixsongs.app and drag the song in. Higher-quality sources give better separations: a WAV or a 320 kbps MP3 beats a 128 kbps file from an old phone.
Step 2: Separate it into stems
Click the AI button on the track. The dialog shows the four stems you’ll get: vocals, drums, bass and music. Stem separation runs on a cloud GPU, so it needs an account and one credit per song. A pack is 15 songs for $2.99, and a failed separation returns its credit automatically. Sign in with Google, then start the separation. It takes about a minute for a typical song.
When it finishes, four stem lanes appear under your track, each with its own fader, M (mute) and S (solo) buttons.
Step 3: Mute the vocals
Click M on the vocals stem. That’s your karaoke track. Press Space and listen through the whole song, especially:
- Choruses, where backing vocals are dense.
- Breakdowns, where the vocal is exposed.
- The very end, where reverb tails can linger.
Step 4: Keep the backing vocals (optional)
Some karaoke singers like a little guidance. Instead of muting the vocals stem, turn its fader down to about −15 dB so the original singer is faint but present. Or draw volume checkpoints on the vocals lane: full level in the choruses for support, silence in the verses.
Step 5: Export the instrumental
Click Export and choose WAV or FLAC for the best quality, or MP3 320 kbps for your phone or a karaoke app. Muted stems are not included in the export.
Fixing common problems
A faint vocal is still audible. This is usually reverb or a doubled vocal that the model put into the “music” stem. Try lowering the music stem slightly during the loudest vocal moments with a couple of checkpoints.
The instrumental sounds thin in places. Some instruments share frequencies with the voice (a lead synth or a sax, for example), and part of them may have gone into the vocal stem. Bring the vocals stem up to about −20 dB to restore body; the voice stays mostly hidden underneath.
Live and very old recordings. Crowd noise, mono recordings and heavy tape saturation are harder for any model. Expect more artefacts.
Quick recap
- Add the song. 2. Click AI. 3. Mute vocals. 4. Export. For the opposite, keeping only the voice, see how to extract an acapella.