Remove text from an image
Drop a photo, screenshot or poster. A text detector finds every caption, subtitle and label; you untick what to keep, and an inpainting model repaints the rest — all in your browser.
Nothing leaves this tab
Workbench
No image handy? Try one of these
What happens to your picture
- 1 · Find the text
PP-OCRv5's detector (4.8 MB, Apache-2.0) marks where letters are — any script, any colour. It does not read the words; nothing is transcribed or logged. On the sample frame this takes about a tenth of a second.
- 2 · You choose
Each find becomes a strip on the picture. Peel off the ones you want to keep, widen the strips if letter outlines are thick, and box anything the detector missed.
- 3 · Repaint the background
LaMa (Samsung Research, Apache-2.0) fills each strip from the pixels around it, in 512 × 512 tiles: about 0.3 s per tile on a GPU with WebGPU, about 11 s on a single CPU thread. Only the strips change; the rest of the file is untouched.
Where it falls short
- Tiny print under about 8 pixels tall and heavily stylised lettering (graffiti tags, script logos) are often not detected. Box them yourself.
- LaMa guesses what was behind the text. On skies, walls and water the guess is usually invisible; a large block of text over a face, a patterned shirt or fine lettering on a sign will be filled with something plausible, not the real thing.
- A page that is text from edge to edge — a receipt, a scanned letter — is split into many tiles and can take many minutes in CPU mode. It works, but it is not what this page is built for.
- One image at a time, up to 12 megapixels. There is no daily cap and no account.
Questions
Is my image uploaded anywhere?
No. The page downloads two model files (the detector and LaMa) and runs them inside this browser tab with WebAssembly or WebGPU. Your picture is read from your disk into memory, processed there and saved back by your browser; no server ever receives it. After the first visit you can disconnect and keep erasing in the same tab.
How does it find the text, and what if it misses some?
PP-OCRv5's text detector produces a heat map of where letters are, and each connected patch becomes one strip. It handles Latin, Chinese, Japanese and Korean, light-on-dark and dark-on-light. It misses very small print and decorative lettering: switch on Box missed text and drag a rectangle over it, and it joins the list like any other strip.
Can I remove the subtitles but keep a logo or a name?
Yes — that is what the list is for. Every piece of text is a separate strip with its own checkbox and a thumbnail of the words, so you can see which is which. Untick the logo, the channel name or the price tag and only the rest gets erased; the unticked areas stay byte-for-byte identical to your original.
How good is it on busy backgrounds?
It depends on what the letters were covering. Subtitles over water, sky, grass or a plain wall usually vanish without a trace. Thick meme captions over fur or a patterned surface come out believable but softer. Big lettering over a face or over other text is repainted with an invention, because the real pixels are gone. Compare with the before/after slider and use Back to the strips to tweak widths or skip a line.
Can it remove watermarks?
We do not offer that. A watermark marks someone else's ownership of the picture, and removing it is not something this page is for. Use it on your own photos and screenshots: captions you added, date stamps from your own camera, subtitles on clips you made.
How is this different from Picsart or cleanup.pictures?
Those send your picture to their servers. cleanup.pictures' free tier exports at 720p, and several other sites cap you at one or two images a day or charge credits. Here detection is automatic like Picsart's, the result is exported at your original resolution with no logo, and there is no daily limit, because the work is done by your own computer.
Why is it slow on some computers?
Without WebGPU the models run on the CPU in a single thread: detection stays under a second, but each 512 × 512 tile of repainting takes around 11 s instead of about 0.3 s. A line of subtitles is usually one or two tiles. Chrome and Edge on desktop have WebGPU; the badge under the Erase button shows which mode you are in.
Whole folders, video frames and bigger images
ConceptCut for Mac erases text and objects across a batch of photos, works on frames pulled from video and is not bound by this page's 12-megapixel limit.
ConceptCut on the Mac App StoreNeighbouring tools
Remove objects from photos — for people, wires and clutter rather than lettering: you paint over what should go.
Image to text — the opposite job — keep the words and copy them out of the picture.