Redact personal information in text
Paste a transcript, an email or a support ticket. Names, phone numbers, addresses and more are found and masked on this device — the text never leaves your browser.
checking your browser The model is fetched once and then kept in this browser's cache.
Redaction workbench
Findings show up on this side.
- The left pane is re-drawn with every hit highlighted by type — hover one to see what the model thought it was and how sure it was.
- The second tab holds the rewritten text, with only the detected characters replaced. Everything else is byte for byte what you pasted.
- Copy it, or download it as a .txt. Nothing is stored — reload the page and it is gone.
Detection settings
How to use it
- Paste, then pick the types. Everything is on by default. Turning off the types you do not care about also makes the scan a little faster, because the types are part of what the model reads.
- Run it and read the highlights. This is the step people skip. The model is good, not perfect — a quick pass over the coloured spans is what turns a guess into a redaction you can defend.
- Take the text, not the page. Copy or download the redacted version. Closing the tab throws away both the original and the result; there is no history and no account.
Keyboard: Ctrl/⌘ + Enter runs the scan from inside the text box.
What it finds, and what it does not
The model is zero-shot: the nine types below are just words handed to it at run time, which is also why it can be honest about where it struggles. Four shapes are matched exactly instead of guessed — email addresses, IPv4 addresses, card numbers that pass the Luhn checksum, and mainland-China mobile numbers — because those have a fixed form and a probability score would only make them less reliable.
| Type | Looks like | Where it slips |
|---|---|---|
| People | Owen Radzik, Dr. M. Feliciano, 张伟 | Bare initials and nicknames are often skipped; a name used as a company name ("Radzik & Sons") usually comes back as an organisation. |
| Email addresses | [email protected] | Well-formed addresses are matched exactly, so they never depend on the model's confidence. Obfuscated forms like "owen (at) fastmail dot example" are not matched at all. |
| Phone numbers | +1 (415) 555-0173, 138 1234 5678 | Mainland-China mobile numbers are matched exactly. Everything else is up to the model, which is most reliable when a label precedes the number ("Phone: …"); a bare run of digits mid-sentence is often read as some other kind of reference number. |
| Postal addresses | 214 Harbor Ridge Road, Apt 6B, Oakland CA 94607 | A span is at most twelve words, so a long address is returned in pieces rather than as one block — read the highlights before you trust it. |
| Dates of birth | 22 April 1979, 1979-04-22 | A date only counts as a birth date when the sentence says so; otherwise it is left alone on purpose. |
| Payment card numbers | 4111 1111 1111 1111 | Any run of 13–19 digits that passes the Luhn checksum is matched exactly. The last four digits alone ("card ending 4242") are not — whether that counts as personal data is your call. |
| ID and account numbers | passport, national ID, member no. NG-884102 | Scores in this class run low (0.43 for the ID number in the Chinese example) because a run of digits has no fixed shape. Keep the threshold down if this class matters to you. |
| Organisations | Northgate Utilities, Brightline Care | Masking every company name tends to make a support ticket unreadable; turn this class off when you need the context to survive. |
| IP addresses | 203.0.113.47 | IPv4 is matched exactly by pattern. IPv6 and MAC addresses are left to the model and are hit and miss. |
Languages: the model was trained on English, French, German, Spanish, Portuguese and Italian. Chinese comes along for the ride and holds up well for names, mobile numbers, emails, IPs and bank card numbers, less so for long addresses. Anything else is a gamble — run your own text through it once before you rely on it.
The model behind the page
- Model
- gliner_multi_pii-v1
- Licence
- Apache-2.0
- First download
- —
- Speed here
- 200 words in ~0.4 s on WebGPU, ~2.9 s on the CPU (single-threaded WebAssembly)
GLiNER is a span model: the types you tick are written into the prompt, the encoder reads them together with your text, and a head scores every run of up to twelve words against every type. It generates nothing — it can only point at characters that are already in your text and say "that is a person", which is exactly the property a redactor needs.
The weights are a quantised ONNX export of gliner_multi_pii-v1 (Apache-2.0), built on an mDeBERTa-v3 encoder (MIT). Your browser fetches them once, caches them, and runs everything locally from then on — including with the network off.
Whole files instead of pasted text
This page takes what you can paste: one document at a time, up to 20,000 characters, plain text in and plain text out. For folders of TXT, MD, CSV and DOCX files — with the formatting kept, a mapping table so you can restore the originals later, and 42 entity presets instead of nine — the same detection lives in our macOS app.
Questions people ask
Does my text get uploaded anywhere?
No. The page has no upload step at all — there is no form post, no API call, no analytics on the text. The only network traffic is the one-time download of the model weights from our CDN, plus the page itself. Open the network panel, run a scan and you will see nothing new appear. That is also why the tool is worth using for text you are not allowed to send to a hosted redaction service.
Which types of personal data, and which languages?
Nine types are wired into the page: people, email addresses, phone numbers, postal addresses, dates of birth, payment card numbers, ID and account numbers, organisations and IP addresses. Each is a plain-language prompt handed to the model at run time, so this is real zero-shot detection rather than a pattern list. Chinese works too: in the Chinese example the name scores 0.99, the mobile number 0.94, the bank card 0.84 and the national ID 0.43 — all measured, not estimated.
How accurate is it? What happens to the things it misses?
There is no honest single number: accuracy depends on your text, your types and your threshold, and anyone quoting you 99% is quoting a benchmark you have never seen. What the page gives you instead is reviewability — every hit is highlighted with its confidence, so a 30-second read tells you what was caught and what was not. For anything with legal weight, treat this as a first pass and read the marked-up view before you hand the text on.
Why do I have to download hundreds of megabytes?
Because the whole model runs on your machine, and a multilingual encoder with a 250,000-token vocabulary is simply that big: 553 MB of weights plus a 16 MB tokenizer, the same file whether you end up on WebGPU or on the CPU. It is fetched once, cached by the browser, and reused on every later visit — the second visit is ready in about a second, offline included. A tool that shipped your text to a server would download nothing at all, which is precisely the trade you are making.
How is this different from a regular expression redactor?
A regex finds shapes: it will happily catch an email or an IPv4 address, and it will never find "Owen", "the Radzik account" or a street address written the way a human types it. This model reads the sentence around a candidate, so names, addresses and dates of birth are on the table — and a date is only flagged as a birth date when the context says so. This page does not make you pick one: email addresses, IPv4 addresses, Luhn-valid card numbers and mainland-China mobile numbers are matched by exact patterns as well, so the shaped things never come down to a confidence score.
Can it do PDF, DOCX or a folder of files?
Not on this page. Here you paste text or drop a .txt file, and you get text back. Our macOS app PIIShield reads TXT, MD, CSV and DOCX in batches, rewrites only the detected characters so the file keeps its structure, and exports a mapping table that can restore the originals. Neither one reads PDF — if that is your input, export it to text first.