Extract text from an image
Paste a screenshot or drop a photo. The characters are read on your own machine, so the picture never leaves this tab.
Add images
Drop images here
⌘V pastes whatever you just screenshotted
or
JPG, PNG, WebP, GIF or BMP · up to 10 images at a time · up to 8 megapixels and 15 MB each
What the model read
Every line is editable — fix a character here and the copy and download buttons take your version.
No text found in that image.
- Zoom in before you screenshot: the detector wants characters at least ten pixels tall.
- Flatten the contrast — pale grey text on a photo background is the usual reason for an empty result.
- Crop away heavy graphics and watermarks; they compete with the real text.
How it works
- The image stays hereIt is decoded into a canvas in this tab. Open the network panel and watch: after the model files, there are no further requests.
- Two small models, one after the otherA detector marks every block of text, then a recogniser reads each block character by character and a CTC decoder turns that into a string.
- You fix and copyLines come back in reading order with a confidence score. Correct the odd character, then copy, or export .txt or .json with the box coordinates.
The models on this page
- Detector
- PP-OCRv5_mobile_det · 4.83 MB
- Recogniser
- PP-OCRv5_mobile_rec · 16.68 MB
- Licence
- Apache-2.0 · PaddlePaddle
- Scripts
- Latin, Simplified & Traditional Chinese, Japanese kana, digits and punctuation — 18,383 characters in one dictionary
PP-OCRv5 mobile is PaddleOCR's small text pipeline, exported to ONNX by the PaddlePaddle team. Both files together are about 21.5 MB, which is why this page can promise the whole thing on a phone. The recogniser handles mixed Chinese and English in the same line without a language switch, because it was trained on one shared dictionary.
Licence text shipped with this page Model card on Hugging Face
Questions
Do my images get uploaded?
No. The only downloads are the two model files (about 21.5 MB, cached after the first visit); the image itself is decoded into a canvas and handed to a worker in this tab. Nothing is sent anywhere, which is the point when the picture is a receipt, an ID, a contract or a medical letter.
Which languages can it read?
One model covers Latin script, Simplified and Traditional Chinese, Japanese kana and the usual punctuation — 18,383 characters in a single dictionary, so a line mixing Chinese and English works without choosing a language. Korean, Thai, Arabic, Cyrillic and Indic scripts are not in this dictionary and will come out as nonsense.
Can it read handwriting?
Only clear, upright handwriting, and not reliably. PP-OCRv5's own benchmark scores about 0.74 on handwritten Chinese against 0.91 on printed Chinese. Screenshots, scans, receipts, slides and signs are what this page is good at; a doctor's note is not.
Why is this faster than an OCR website, and where is it slower?
There is no upload and no queue, so a screenshot is read in about 0.4 seconds on WebGPU or 2–3 seconds on the CPU path (both measured on an Apple M2 Max), however busy the internet is. The slow part is the first visit: 21.5 MB of model files have to arrive before anything happens. After that the page also works offline.
Can it read a PDF?
Not directly — this page takes image files. For a single page, screenshot it or export it as PNG and drop that in. Multi-page PDF handling, batches beyond 10 images and the larger server-grade models are what the paid tier is for.
Does it keep tables, columns and layout?
It keeps reading order, not structure. Lines are grouped top to bottom and left to right, with a blank line where there is a visible gap, so a receipt or a form comes out readable — but a wide table loses its columns, and two-column pages interleave. Each block's four corner coordinates are in the .json export if you need to rebuild the layout yourself.