Remove vocals from a song — a free vocal remover that runs in your browser

Drop a take in. Vocals and instrumental are pulled apart on this machine — nothing is uploaded, nothing waits in a queue.

Heads up on a phone: the model is a 66.8 MB download and a 3-minute song is 32 passes through it. Clips under a minute are fine; full songs want a laptop or desktop.

Load a take

…or drop the file anywhere on this page.

mp3, m4a, wav, flac, ogg, opus, mp4 — whatever your browser can decode. One file at a time, up to 10 minutes and 50 MB.

20 seconds we made for this page: band, vocal line, band again. Runs end to end, before you commit to the 66.8 MB download.

What actually happens when you press Split

  1. Your browser decodes the audioThe file is read with the same decoder your browser uses for playback, resampled to 44,100 Hz stereo, and kept in this tab. It is never sent anywhere.
  2. The take is cut into 5.75-second passesEach pass is turned into a spectrogram with a 7680-point FFT hopping 1024 samples — exactly the analysis the model was trained on, with 87 ms of overlap on each side so the seams stay inaudible.
  3. The model keeps only the voiceUVR-MDX-NET Voc_FT reads the 3072 lowest frequency bins and returns the part it believes is singing. An inverse FFT turns that back into sound.
  4. The instrumental is the leftoverVocals are subtracted from your original sample by sample, so the two stems add back up to exactly what you loaded — no re-encoding, no loudness tricks.

Timings from this machine: on WebGPU a pass takes about 0.3 s, so a 3-minute song is done in roughly 20 seconds. In CPU mode everything runs on one WebAssembly thread and a pass takes about 19 s, so the same song took about 10 minutes. Nothing is cached server-side, so the second song is as fast as the first.

The model on this page

Model
UVR-MDX-NET Voc_FT (ONNX)
One-time download
66.8 MB, fp32
Licence
MIT — read the full text
Speed here
≈0.3 s / 5.75 s pass on WebGPU

Weights by Anjok07 and aufr33 of the Ultimate Vocal Remover project, MDX-Net architecture by Kuielab. The UVR authors licence their models under MIT and ask that anyone shipping them gives credit — consider this page's credit paid, and the exact wording is in the licence file. Inference runs on ONNX Runtime Web (MIT).

Questions people ask before dropping a file in

Is my music uploaded anywhere?

No. The only thing that crosses the network is the 66.8 MB model file, and it travels towards you, not away. Open DevTools → Network before you press Split and watch: after the model is in the cache there are no further requests. That is also why there is no queue and no minute counter.

How does this compare with LALAL.AI, vocalremover.org and the rest?

They upload your file and run bigger models on their servers, then meter you by the minute — LALAL.AI's free tier is ten minutes total, for instance. This page runs a 66.8 MB model on your hardware: unlimited, free and private, but your laptop does the work and the ceiling is lower on dense mixes. For clean pop, rock and singer-songwriter material the difference is hard to hear; on crowded masters with heavy reverb a server-sized model still wins.

Can I get the drums and bass as separate tracks?

Not on this page — you get two stems, vocals and instrumental. Four-stem separation needs a different model (htdemucs, 165 MB, roughly three times the computation per second of audio). It is on the list for a paid tier here; it is not hiding behind a button today.

Why is CPU mode so much slower than WebGPU?

Same maths, different hardware. One 5.75-second pass takes about 0.3 s on a GPU through WebGPU and about 19 s on a single WebAssembly thread; for a whole 3-minute song that is roughly 20 seconds against 10 minutes. Chrome and Edge on a recent machine get WebGPU; Firefox and older Safari fall back to the CPU path automatically, and the status line above tells you which one you are on.

Can I drop in a video file?

Yes, if your browser can decode it — .mp4 and .webm usually work. Only the audio track is read, and what comes back out is audio: two WAV files, no video. If the file is DRM-protected nothing can be decoded, and you will get a message saying so.

What am I allowed to do with the stems?

That is between you and whoever owns the recording. Separating a track you own, licensed, or recorded yourself — a demo, a rehearsal tape, a client's mixdown — is routine studio work. Pulling an a cappella out of a commercial release and publishing it is not, and running it in your browser rather than on someone's server does not change that. Nothing here checks; you are responsible.

What are the limits, and what do I actually get back?

One file at a time, up to 10 minutes and 50 MB. Out come two 44.1 kHz 16-bit stereo WAV files — vocals and instrumental — that add back up to your original sample for sample. They live in this tab only: reload and they are gone, because nothing is stored anywhere.