Flatten a photo of a document
Photograph a page, even a curved one. The flattening runs on your device; the file never uploads.
Your pages
Drag the four yellow dots onto the corners of the paper. The result is cut to that shape and its proportions are taken from the edges you mark, so a long receipt comes out long.
How it works
Turn it the right way up
A small classifier (PP-LCNet, 6.8 MB) looks at the page and decides whether it's sideways or upside down, then rotates it before anything else happens.
Predict how the paper is bent
UVDoc shrinks the photo to 712 × 488 and predicts a 45 × 31 grid that says where every part of the flat page ended up in your photo — the bend of the spine, the fold, the tilt of the camera.
Pull every pixel back
That grid is stretched to your full resolution and each output pixel is read from where the grid points. The result keeps your photo's pixel size; trim and finish it, then save.
Why a four-corner crop can't fix a curved page
Most free scanners online, and the auto-crop in many phone apps, find the four corners of the paper and apply one perspective transform. That is exactly right for a flat sheet photographed at an angle. A book page near the spine is not flat: it curls away from the camera, so its lines of text are arcs, and no single perspective transform can turn an arc into a straight line.
UVDoc was trained on photos of creased, curled and folded paper paired with their flat originals, so it predicts a different correction for every part of the page. Our crumpled-receipt example shows it on folds; the book-page example shows it at the spine.
Taking a photo that flattens well
- Get the whole sheet in frame with a strip of table showing on every side. The model needs to see the page edge to know where the paper ends.
- Shoot from above rather than at a steep angle. The model handles tilt, but text that is tiny in the photo stays tiny after flattening.
- One page per photo. With an open book, photograph the left and right pages separately; a two-page spread is treated as one sheet.
- Light from the side leaves a shadow in the gutter. “Lift shadows” evens it out; “Black & white” is for text you want to print or fax.
The model and its licence
Flattening is done by UVDoc by Floor Verhoeven, Tanguy Magne and Olga Sorkine-Hornung (ETH Zurich, SIGGRAPH Asia 2023, MIT licence), in the ONNX export published by PaddlePaddle as PaddlePaddle/UVDoc_onnx (Apache-2.0, 31.7 MB). Orientation comes from PP-LCNet_x1_0_doc_ori (Apache-2.0, 6.8 MB).
Both run through ONNX Runtime Web inside this tab — on your graphics card via WebGPU when the browser offers it, otherwise on the CPU with WebAssembly. Full licence texts: LICENSE-model.txt.
Measured on our test machine (Apple M2 Max, Chrome): a 1600 × 1200 photo flattens in 0.06 s with WebGPU and 0.95 s on the CPU (WebAssembly runs on a single thread here); a 12-megapixel phone photo takes 0.2 s and 1.2 s. An ordinary laptop's CPU will take longer.
Questions
Is my photo uploaded anywhere?
No. The page downloads the two model files and then does all the work in your browser tab. Your photo, the flattened image and the PDF are made on your device and never sent to a server — you can switch off Wi-Fi after the model has loaded and keep going in the same tab.
How is this different from my phone's scanner or a four-corner crop tool?
Those find the four corners of the page and square them up, which fixes the camera angle but leaves a curved page curved. This tool predicts how the paper itself is bent and undoes that, so the lines of text near a book's spine or across a fold come out straight. The two diagrams above show the difference.
Why does the first run download 37 MB?
That is the two neural networks: UVDoc (31.7 MB) and the orientation classifier (6.8 MB), unquantised because smaller versions don't exist for this model. They are stored in your browser's cache, so the next visit starts in about a second without downloading again.
My browser doesn't have WebGPU. Will it work?
Yes. Without WebGPU the same model runs on the CPU with single-threaded WebAssembly. On our M2 Max that is about 1 s instead of 0.06 s for a typical photo, and slower laptops need longer. Nothing about the result changes — we compared the two paths pixel by pixel and they differ by at most 1 in 255.
Can I do several pages and get one PDF?
Yes. Add up to 5 photos per batch, reorder them with the arrows under each thumbnail, then press “Save all as PDF”. Each page keeps its own proportions and the finish you picked. The PDF is written by this page itself, not a web service.
Can it pull the text out of the page?
No, this tool only makes the image flat. To turn the flattened page into editable text, save it and open it in our image to text tool, which also runs locally — flattened pages are read noticeably better than curved ones.
What kind of photos does it get wrong?
Pages cut off by the frame (it can't flatten what it can't see), two-page spreads (they are flattened as one sheet), and very small text on a distant page. It also keeps the photo's pixel size, so a receipt comes out as wide as the photo was — use “Trim corners” to cut it to the paper.