Remove background from video — free, with sound, in your browser
Drop a clip. Every frame is processed on your device, and the audio comes along with it.
This works on a phone, but a 30-second clip is several times faster in a desktop browser.
Your clip and export settings
Drop a video clip here
or paste it with Ctrl+V / ⌘V
mp4 · mov · webm — up to 30 s, 720p and 200 MB. Longer or bigger clips are trimmed or scaled, not refused.
00:00
person foundno person in frameClick a bar or use ← → to step through frames
What happens to your clip
- 1 · Decode
Your browser's own video decoder (WebCodecs) unpacks the clip frame by frame. Only one frame is in memory at a time, so a 30-second clip doesn't need gigabytes.
- 2 · Matte
Each frame is shrunk to 320–512 px on the short side and handed to MODNet, a portrait-matting network, which returns a soft alpha mask. That mask is scaled back up to your frame and multiplied in.
- 3 · Encode
The cut-out frames are re-encoded — VP9 with alpha for WebM, H.264 for MP4 — and the audio packets from your file are copied into the new container.
How long it takes
Measured on the 6-second 720p example (180 frames) with Standard edge detail. Your times scale with clip length.
| Machine | Model per frame | 6 s example | 30 s clip (estimate) |
|---|---|---|---|
| Apple M2 Max, WebGPU | 55 ms | 11 s | about 55 s |
| Same machine, CPU only (single thread) | 660 ms | about 2 min 10 s | about 11 min |
An older or low-power laptop without WebGPU will be slower still than the CPU row. If it's slow, trim the clip first — every second you cut is 30 frames the model doesn't have to look at.
The model
Matting is done by MODNet (Ke et al., AAAI 2022), a network trained to separate people from their surroundings, in the ONNX export by Xenova/modnet. With WebGPU the 25.9 MB full-precision weights run on your graphics card; without it the 13.0 MB half-precision file runs on the CPU. We tested the 6.6 MB 8-bit versions too: on a light-blue shirt in front of a white wall they erased the whole person, so they are not used.
MODNet is released under the Apache License 2.0, which allows commercial use. The full text ships with this page: LICENSE-model.txt. Video decoding and encoding use Mediabunny (MPL-2.0).
Questions
Is my video uploaded anywhere?
No. The page downloads the model once, then decoding, matting and encoding all happen inside this browser tab. The server only ever sends files to you; there is no upload endpoint. You can disconnect after the model has loaded and keep processing clips in the same tab.
Why does my transparent WebM play with a black background?
The transparency is stored as a second VP9 stream beside the picture. Chrome, Edge and Firefox read it; QuickTime Player, Windows Photos, phone galleries and many desktop editors ignore it and draw black. The file isn't broken. If your editor doesn't take WebM alpha, export PNG frames (an image sequence every editor imports) or the chroma-green MP4 and key it there.
Is the sound kept?
Yes. For MP4 the audio packets are copied from your file untouched, so there is no quality loss. WebM can't carry AAC, so for transparent WebM the audio is re-encoded to Opus. The PNG export puts the audio in the zip as a separate file. Some free web tools drop the audio from their free exports; this one doesn't.
Why does the first run download 13–26 MB?
That's the MODNet model: 25.9 MB in full precision for WebGPU, 13.0 MB in half precision for the CPU path. It is stored in your browser's cache, so the second visit loads it in about a second and downloads nothing.
Does it work without WebGPU, and how long does it take?
Yes — without WebGPU the model runs on the CPU through WebAssembly, on a single thread. On an Apple M2 Max that is roughly 660 ms per frame at Standard detail, so a 30-second clip at 30 fps (900 frames) takes around ten minutes; an older laptop takes longer. Quick edge detail (about 410 ms per frame) cuts the time by roughly a third.
How long a clip can I process, and why is there a limit?
The free limit is 30 seconds at up to 720p and 200 MB. Longer clips are trimmed to their first 30 seconds and bigger ones are scaled to fit 1280×720, so you still get a result. The limit exists because everything runs on your machine: a minute of 1080p is 1,800 full-size frames, which is slow on a CPU and heavy on memory in a browser tab.
What kind of footage works best?
MODNet is trained on people, so it's at its best on talking heads, presenters, dancers and anyone facing the camera. It doesn't cut out pets, cars or products. Hair against a background of the same colour, fast motion blur and people who are very small in the frame give softer or patchier edges — the frame timeline shows which moments had no person detected.
Only need one still frame? Remove background from image does a single photo at full resolution.