Porting UVR's MDX-Net vocal remover to the browser: the model was easy, the STFT wasn't
Our vocal remover splits a song into a vocal track and an instrumental track entirely inside the browser tab. It uses UVR-MDX-NET Voc_FT, one of the models from the Ultimate Vocal Remover project (by Anjok07 and aufr33, MIT-licensed). The model is a 66.8 MB fp32 ONNX file, and ONNX Runtime Web runs it on WebGPU or on WebAssembly with no changes.
Getting the model to run is the easy part. Getting it to produce vocals rather than noise is where the work is. This post covers that work: the signal processing that the ONNX file doesn't contain, how we checked it sample by sample against PyTorch, and a check we wrote that turned out to prove nothing.
The model only speaks spectrogram
If you open UVR-MDX-NET-Voc_FT.onnx, the graph is a plain convolutional U-Net: 66 Relu, 40 Conv, 27 BatchNormalization, 22 MatMul, 5 ConvTranspose, opset 13. There's no STFT op anywhere. The input and output are both a tensor of shape [batch, 4, 3072, 256]:
- 4 planes: real and imaginary parts for the left channel, then for the right;
- 3072 frequency bins: the lowest 3072 of the 3841 bins a 7680-point FFT produces;
- 256 time frames, with a hop of 1024 samples at 44.1 kHz, which is about 5.9 seconds of audio per call.
So the browser has to do what UVR's Python code does around the model: cut the audio into chunks, run a short-time Fourier transform with exactly the settings the model was trained with, pass the complex spectrogram through the network, run the inverse transform, and stitch the chunks back together. Transformers.js has no audio-to-audio pipeline to lean on, so this became a custom engine: about 340 lines of JavaScript that run inside a Web Worker.
Don't guess the parameters: UVR looks them up by hash
An MDX-Net checkpoint carries four numbers that aren't stored in the ONNX file: n_fft, dim_f, dim_t and a gain called compensate. Blog posts and forum threads quote different values for different models, and they're easy to mix up. Get one wrong and the network still produces a spectrogram, just not one that means anything.
UVR resolves this with a table, model_data.json in its application_data repository (59 entries when we fetched it). The key is not the file name, and it isn't the hash of the whole file either. It's the MD5 of the last 10,000 × 1024 bytes of the weights:
with open('UVR-MDX-NET-Voc_FT.onnx', 'rb') as f:
f.seek(-10000 * 1024, 2)
print(hashlib.md5(f.read()).hexdigest())
# 77d07b2667ddf05b9e3175941b4454a0
That hash maps to n_fft 7680, dim_f 3072, dim_t 2^8, compensate 1.021, primary stem "Vocals". The full-file MD5 has no entry. It's also worth knowing that neighbouring entries in the table are close but not the same. One has the same FFT size and bin count but a compensate of 1.045 and is an instrumental model. If you copy a parameter set from the wrong line, you get a slightly wrong gain and never notice.
7680 is not a power of two
7680 = 29 · 3 · 5. Most small JavaScript FFT libraries are radix-2 only, so they can't handle it. We wrote a recursive mixed-radix Cooley–Tukey transform with radix-4, radix-2, radix-3 and radix-5 stages. The first step factors the size:
function radicesOf(n) {
const out = [];
let m = n;
for (const r of [4, 2, 3, 5]) while (m % r === 0) { out.push(r); m /= r; }
if (m !== 1) throw new Error(`fft size ${n} has a prime factor > 5`);
return out;
}
Radix-2 and radix-4 get hand-written butterflies. The odd radices use a small generic DFT, which is fine because they only show up twice in the factorisation. Twiddle factors are precomputed once in Float64Arrays. We also used the standard trick of computing a length-n real FFT as a length-n/2 complex FFT plus a post-processing pass, so each 7680-point real transform is really a 3840-point complex one. Everything runs in float64. The only float32 data is the tensor handed to ONNX Runtime.
Matching torch.stft, detail by detail
The FFT was the easy half. The hard half is matching what torch.stft and torch.istft do by default, because the network learned from exactly those spectrograms. Three details matter:
- Periodic Hann window:
0.5 − 0.5·cos(2πi/N), not the symmetric variant that divides byN−1. - Centre padding by reflection:
center=Truepadsn_fft/2samples on each side, mirrored without repeating the edge sample:function reflectPad(x, out) { const n = x.length; out.set(x, TRIM); for (let j = 0; j < TRIM; j++) out[j] = x[TRIM - j]; for (let j = 0; j < TRIM; j++) out[TRIM + n + j] = x[n - 2 - j]; } - Inverse normalisation by the squared-window envelope:
torch.istftoverlap-adds the windowed frames and divides by the overlap-added sum ofwindow². The envelope depends only on the geometry, so we compute it once and guard the division at the edges, where the envelope approaches zero.
None of this is new. It's what the PyTorch source does. But every one of these is a place where "roughly right" produces a result that sounds roughly right and isn't, so we treated them as things to copy exactly, not to redesign.
Proving it: a PyTorch reference, compared sample by sample
Before building any UI, we wrote two programs that process the same 20-second WAV:
reference.py:torch.stft→ ONNX Runtime (Python, CPU) →torch.istft, following UVR's own chunking;compare.mjs: ourengine.js→ ONNX Runtime (Node, CPU), dumping the raw float32 spectrogram and waveform.
Then a small script compared them:
| What | Correlation | Relative error | Max |Δ| |
|---|---|---|---|
| Input spectrogram (4 × 3072 × 256 per chunk) | 1.000000 | −136.8 dB | 6.1e-5 (peak 533) |
| Vocal waveform | 1.000000 | −132.4 dB | < 1e-6 |
A −132 dB relative error is float32 rounding, not a real difference. The engine was correct on its first full run. We put that down to looking up the parameters before writing any code, not to cleverness.
A caveat about the test material: the 20-second clip is synthetic. The accompaniment is generated with numpy, and the "singer" is macOS text-to-speech with pitch shift and vibrato, entering at 5 s and stopping at 13.55 s. That makes it good for checking that two implementations agree, and for measuring leakage in the sections where the vocal is silent. It says nothing about how the model sounds on your favourite record. For that, try it on a real song.
The check that couldn't fail
Following UVR, the instrumental is computed as mix − vocals, not produced by the model. Our first acceptance spec included "vocals + instrumental must sum back to the original within −40 dBFS, to validate overlap-add". The page reports −240 dBFS, which looks superb.
It's also meaningless. vocals + (mix − vocals) − mix is exactly zero by construction, and −240 dB is just 20·log10 of the epsilon we add to avoid log(0). The check can't fail whatever the STFT does.
The obvious fix doesn't work either. A plain round trip (STFT → iSTFT, no model) against the input only reached about 15 dB of separation between residual and signal. The reason is legitimate: the model only sees bins 0–3071, which cut off at 3072 × 44100 / 7680 = 17,640 Hz, and we zero everything above that before the inverse transform. So a round trip is supposed to change the signal.
What does work is idempotence. A correct pipeline is a projection: run the round trip once and it removes the content above 17.6 kHz; run it again on its own output and nothing should change. A wrong window, a missing normalisation or an off-by-one in the padding shows up at once. In Chromium the second pass differs from the first by −151 dB, and in WebKit by −127 dB. That's the number in our acceptance report that actually tests the overlap-add. The −240 stays in the report as a reminder of what a vacuous test looks like.
When the GPU gets fast, the JavaScript shows
These are the raw ONNX Runtime numbers for one chunk on an Apple M2 Max in headless Chromium (ORT Web 1.26-dev):
| Backend | Model, per chunk | Whole chunk incl. STFT/iSTFT |
|---|---|---|
| WebGPU | ≈ 0.31 s | ≈ 0.54 s |
| WebAssembly, 4 threads | ≈ 4.8 s | ≈ 5.1 s |
On the CPU path the DSP is a rounding error. On WebGPU, going by the difference between the two columns, the JavaScript transforms take roughly 0.23 s per chunk, about 43% of the total. That's 1024 complex 3840-point FFTs per chunk: two channels, 256 frames, forward and inverse. End to end, the 20-second sample takes 2.2 s on WebGPU and 19.8 s on WebAssembly. A 3-minute-4-second song takes 16.9 s and 157.5 s, with the JavaScript heap peaking at 202 MB.
We haven't optimised the transforms yet. The obvious next steps are splitting the two channels across workers, moving the FFT to WebAssembly with SIMD, or doing the STFT on the GPU. The general point is worth writing down anyway: for models that take features rather than raw data, pre- and post-processing are the bottleneck to watch. A GPU that speeds up the network by 15× leaves whatever runs around it as the next thing in line.
Smaller things we'd tell someone doing the same
- Chunk the work in the page, not in the engine. We call the engine once per 5.9-second chunk. That gives an honest progress bar and a Stop button that works, and it keeps memory flat. The cost is about 2 MB of structured-clone copying per chunk, which is negligible next to the rest.
- Context-and-trim instead of crossfades. Each chunk is read with
n_fft/2samples of context on both sides and only the middle is kept, which is exactly how UVR tiles the song. There's no crossfade to tune. - onnxruntime-node has no
wasmexecution provider. Our comparison harness first failed with "backend not found" because it reused the browser configuration. A dev-onlyproviders: ['cpu']override fixed it. - fp32 only. UVR ships these weights as fp32 and we didn't make our own fp16 or int8 export. 66.8 MB was already within budget, and a new quantisation would have meant re-validating the one thing we'd just proved.
Try it
The vocal remover runs the engine described here. It takes a song up to 10 minutes long and gives you two 16-bit 44.1 kHz WAV files, and the audio never leaves your machine. The tool page credits Ultimate Vocal Remover and includes the full licence text. If you're doing something similar and want to compare notes on the STFT, write to us.
Tools mentioned in this article: Remove vocals from a song — a free vocal remover that runs in your browser