Create karaoke tracks easily: split any song (YouTube, social links or MP3/WAV) into up to 6 stems, sing along with synced lyrics, record your own takes and hear the song in your own voice with AI voice cloning. Fast, high-precision vocal isolation with cutting-edge AI.
The vocal remover separates a finished mix back into its parts. Give it a song from YouTube, SoundCloud, TikTok or Facebook, or upload an MP3 or WAV, and a neural network trained on thousands of multitrack recordings predicts which parts of the sound belong to the voice, the drums, the bass, the piano, the guitar and the remaining instruments. You get each part as its own file: a clean instrumental for karaoke, an acapella for a remix, or a drum-only track for practice.
And it does not stop at the split. The result opens in a karaoke studio: the lyrics light up word by word as the song plays, you can record your own takes over the band, and AI voice cloning sings the song back in your own voice. Everything runs on Mazmazika's servers — the Ultra and High Quality engines on GPUs — so it works on a phone as well as on a laptop, with nothing to install.
Read live from the current plan settings, so this section always matches what the tool enforces today.
See the plans and current prices
Compared with a classic phase-cancellation "vocal cut" or a desktop stem splitter, the online remover needs no install and no GPU of your own, the AI separation keeps the reverb tails and harmonies that a centre-channel trick would erase, and the result is a complete karaoke studio rather than a folder of files.
The studio shows every stem as a track: horizontal lanes with a waveform you can click to jump, or classic vertical channel strips. Each track has rotary volume and pan knobs, mute and solo, stereo level meters and its own effect rack — low cut, three-band EQ, compressor, warmth, transpose, double, flanger, echo, reverb and limiter — with one-tap presets such as Polish, Warm, Air, Hall, Radio and Slapback.
Recording is built in: press Record and a three-beat count-in starts the band, then your voice is captured clean over the stems. Every take can become a track of its own — Take 1, Take 2 and so on — and you can name each one (up to 10 characters) while you sing, so you always know which take is which. You can also upload a take you recorded elsewhere. Downloads render exactly what you hear: one track, the whole mix or a ZIP of every track, with or without the effects.
Switch on Synced lyrics before processing, or add them later from inside the studio — only the separated vocal is read again, the song is not processed a second time. The words are transcribed from the isolated voice and each one is timed to the music, so the karaoke screen lights them up as they are sung.
Word-by-word timing is available in 16 languages; songs in every other language still get their lyrics, timed line by line. Right-to-left scripts such as Arabic and Persian are shown in their correct order.
"Add a track in your voice" sings the song's lead vocal again in your voice. Record or upload 10 to 30 seconds of yourself singing or speaking, confirm it is your own voice, and a singing voice conversion model re-sings the separated vocal with your voice while keeping the original melody, timing and words. The new track arrives beside the original: the original lead is muted so you hear yourself, and one tap on M brings it back to compare.
Two voice cloning engines are available: Seed-VC, the faster one, and SoulX-Singer, which holds the melody very closely and takes about twice as long. You choose the octave (whole octaves only, so your voice stays in key with the band) and the number of diffusion steps — more steps give more detail and take proportionally longer. Each clone becomes a new track, up to eight voice tracks per song, so you can compare the engines side by side.
On well-produced pop, rock and hip-hop the instrumental is clean enough for karaoke and live use. Very reverberant recordings, live bootlegs and songs where the voice is doubled by a synth can leave faint traces. The High Quality and Ultra engines reduce those artifacts noticeably.
MP3, WAV and M4A files, plus links from YouTube, SoundCloud, Facebook and TikTok. Output stems are MP3 at 128 kbps on the Standard engine and 320 kbps on High Quality and Ultra.
A stem is one group of instruments exported as its own audio file: vocals, drums, bass, piano, guitar, and "other" for synths and everything left. Two-stem mode gives you just voice and instrumental.
Word-by-word timing in 16 languages: Arabic, Chinese, Dutch, English, Finnish, French, German, Greek, Hungarian, Italian, Japanese, Persian, Polish, Portuguese, Russian and Spanish. Songs in any other language still get their lyrics, timed line by line, so the right line lights up while it is sung.
Any. Voice cloning for singing changes only the voice: the melody, the timing and the words come from the original singer, so a song in Arabic, English or Japanese is re-sung in your voice in the same language. The two engines learned from different data — SoulX-Singer from Mandarin, English and Cantonese singing, Seed-VC with a multilingual speech encoder — so if the words of one song come out unclear, try the other engine.
Yes. The sample is stored in your account only after you confirm it is your own voice or that you have the speaker's permission. It is used only to make your voice tracks, and you can replace or delete it at any time. The copy sent to the GPU for a job is deleted with that job.
As many as you like, one after another. With "Each take becomes its own track" switched on, every take is a separate track you can name, mix and download; a song holds up to eight voice tracks (takes and voice clones together). Switch the option off and your takes add up in a single "You" track instead.
The separation itself is yours, but the underlying recording keeps its copyright. Karaoke at home, practice and private study are fine; releasing a remix or performing publicly needs a licence from the rights holder. A voice clone may only be made of your own voice, or with the speaker's permission.
Each engine has a per-run track length that depends on your plan; the exact minutes are listed under Free vs Pro above and on the tool itself. The limit protects the shared queue so results stay fast for everyone.
Yes. The processing happens on our servers, so the page only needs a browser. The karaoke screen, recording and the mixer all work with touch, and downloads land in your phone's files app.