Background removal in the browser: 4 open models tested on WebGPU and WebAssembly

Blog · · SwordFishSoft

If you want to remove backgrounds without uploading anything, the model has to run on the user's machine. Several open models do this well in Python. Far fewer come as an ONNX file that runs correctly in a browser, fits a sensible download and has a licence that allows commercial use.

We ran the four candidates that got through those filters on the same photos, on WebGPU and on WebAssembly, in the same headless Chrome. This post has the numbers, the cut-outs side by side, and the reasons the other well-known options didn't make the list. We first made this comparison on a single photo when choosing the model for our background remover. This is the fuller re-run.

The short version

Model (Hugging Face repo)LicenceDownloadWebGPU / imageCPU (WASM) / imageGood atBreaks on
BiRefNet-lite, 512² export
studioludens/birefnet-lite-512
MIT98 MB (fp16)
192 MB (fp32, CPU)
1.1 s1.6 sObjects, pets, portraits: the most consistent of the fourVery fine strands (whiskers). Can drop a secondary subject cut off by the frame
ISNet general, int8
xrds/isnet-general-onnx-int8
MIT44 MB0.32 s2.2 sSpeed on GPU; keeps everything that might be foregroundLeaves translucent haze and keeps tables or props you didn't want
BEN2
onnx-community/BEN2-ONNX
MIT219 MB (fp16 only)0.98 s14.1 sClean single subjectsOur group shot: holes and fragments. Too slow on CPU
MODNet
Xenova/modnet
Apache-2.026 MB0.14 s0.44 sPeople and animals, very fastAnything that isn't a person or animal: the teapot fell apart

Times are the median per image across four photos (the three below plus a group shot), after one warm-up image. They include decoding the JPEG and the pipeline's own pre- and post-processing. Machine: Apple M2 Max, headless Chromium 153, Transformers.js 4.2.0 with onnxruntime-web 1.26.0-dev, WebAssembly with 12 threads. Measured on 2026-09-22.

Side by side

A grid of three photos (a black teapot, a corgi, a woman holding white flowers) and the cut-outs from four models on a checkerboard. MODNet breaks the teapot into fragments; the other results look broadly similar at this size.
WebGPU outputs on a checkerboard. The photos are CC0: "Black Ceramic Teapot" and "Female in Sunglasses Holding White Flowers" by Image Catalog, "Office dog portraits" by Lottie's pets & stuff, all via Flickr/Openverse, resized to 1024 px.

At thumbnail size three of the four look fine on all three photos, and MODNet's teapot is the obvious failure. The differences show up when you zoom in:

Close-up crops of the corgi's ear and the flowers against the woman's arm for each model. None of the models keeps the whiskers; ISNet leaves a translucent patch beside the arm.
Crops at 75% of full resolution: the corgi's ear and whiskers, and the bouquet against the arm.

The group photo

We also ran a fourth, harder photo: people around a table of books outdoors, with one person partly cut off at the left edge. We aren't showing it because we couldn't confirm its licence, but the numbers tell the story. The foreground fraction varied from 18.7% (BEN2) to 39.8% (ISNet) on the same image:

"What counts as the subject" is a judgement call. No model gets it right for every purpose, which is one reason the tool you use should make it easy to see the mask before you download.

WebGPU vs WebAssembly: same answer, different speed

For each model, the two backends gave the same foreground fraction to within 0.2 percentage points on every photo. That's reassuring, because it isn't always true (we've written about models that silently return garbage on WebGPU). Speed is another matter:

When this was written, our tool capped WebAssembly at 4 threads so the rest of the page stayed responsive. In our acceptance run at that setting, BiRefNet-lite took 1.8 s for a 1024×683 photo on the CPU and 1.0 s on WebGPU. Update, 25 September 2026: the site no longer sends the cross-origin isolation headers, so the CPU path now runs on one thread. The same photo takes 5.3 s on the CPU; WebGPU is unchanged.

Working around a 512×512 mask

A model that predicts its mask at 512×512 doesn't have to give you a 512-pixel result. Here's what our tool does with BiRefNet-lite's output. None of it is specific to that model.

None of this makes the mask sharper than the model predicted, and whiskers stay lost. What it does is keep the rest of the image at full quality, so the model's resolution only affects the edge.

What didn't make the list, and why

Which one to use

How to reproduce

Each (model, backend) pair is loaded once through Transformers.js's background-removal pipeline. It's warmed up on one image, then timed on each photo, and the RGBA result is saved. The foreground fraction is the share of pixels with alpha above 127. The WebGPU and WebAssembly runs share one browser profile so each file downloads once. The script is about 60 lines of Playwright plus a small HTML page, and the only non-default setting is graphOptimizationLevel: 'basic' for BEN2's fp16 graph on WebAssembly. Write to us if you'd like a copy.

Tools mentioned in this article: Remove background from image — free, full resolution, in your browser