Home / Document redaction software
Document Redaction Software
Three ways to find what to remove — keyword, pattern, or AI detection — across PDF, Word, Excel, PowerPoint, Visio, OneNote, CAD, email and HTML. Including the photograph on page four that text-only tools cannot see.
Why Teams Move
4M pages
Redacted a year on a single deployment
6
Script families read, from Latin to Perso-Arabic
63
Document formats in, no conversion step
What Makes Documents Hard
Scans are rarely clean
Photocopies arrive skewed, rotated, faint, or photographed on a phone at an angle. Text with no text layer behind it. OCR has to cope with the document you were sent, not the one you wish you had.
Documents are not just text
A police report carries a booking photo. A claims file carries a vehicle image with a plate in it. A benefits application carries a scanned ID. A text-only tool redacts the character runs and leaves the photograph — and the redaction looks complete.
One page, several types
Printed text, handwriting, a signature, an embedded image and a table, all on the same sheet, each needing different handling.
One column, 400 rows
Spreadsheets need redaction by column and by row, not cell by cell.
Not every document is English
Each text region is classified by script and routed to the engine that reads it — Latin including Welsh, Perso-Arabic covering Arabic, Urdu, Sindhi, Dari and Pashto, plus Devanagari, Cyrillic, Chinese and Japanese, and Korean. Routing is per region, so a mixed-script page is handled region by region.
The same form, every week
It arrives on a schedule and gets marked up by hand every time.
Formats nobody planned for
, and metadata that survives the redaction unless it is stripped on export.
How to Find What to Redact
Keyword
Specific words or phrases, wherever they appear
Pattern
Regular expressions and predefined templates such as SSN or credit card
AI-assisted
PII, text recognised through OCR, and objects such as faces and signatures
Templates for recurring types
Where documents share a layout, the regions carrying personal data sit in the same places. A template records those regions and the classes to detect, then applies to every document of that type, including across a bulk batch. That is what makes a queued batch consistent rather than dependent on which operator ran it.
What You Can Redact
Printed text
PII, PHI, CCI and financial data in the document body
Scanned text
Pages with no text layer, recognised through OCR
Handwriting
Handwritten entries, through intelligent character recognition
Signatures
Detected as a class of their own
Embedded images
Faces, plates and vehicles inside photographs on the page
Table columns and rows
A whole field of values in one action
Non-Latin scripts
Perso-Arabic, Devanagari, Cyrillic, CJK and Korean
Keywords and patterns
Specific phrases, or regex and templates such as SSN
Document metadata
Attributes carried in the file, stripped on export
Proof
VIDIZMO
86.1%
Idox plc
32.8%
ReadyRedact
30.7%
The Department for Work and Pensions in the UK runs the largest documented deployment: roughly 5,000 users processing 4 million pages a year on a single installation. That answers whether the architecture holds at document volume rather than in a demo.
It also produced the most useful competitive evidence available, because the scoring above was done by the buyer against public-sector requirements — not by a vendor writing its own comparison table.
Compared With Adobe Acrobat
| Acrobat | Redactor |
|---|---|
| One file at a time, one person at a time | Bulk batches, queued and worked overnight |
| Text only; an embedded photo is opaque to it | Page rasterised, so faces and plates inside PDFs are detected |
| Mark up each recurring form by hand, every time | Templates apply the pattern to every document of that type |
| Cell by cell on a spreadsheet | By column and by row |
| Needs a text layer before it can help | OCR and handwriting recognition built in |
| No exemption codes, no coverage report | Statutory code on every redaction, printed on the output |
Where It Runs
Shared SaaS, dedicated cloud, your own private cloud, on-premises, hybrid, or fully air-gapped.
FAQ
Document Redaction Software questions, answered
Is this different from using Adobe Acrobat?
Acrobat redacts one file at a time and works on the text layer. It cannot see a photograph embedded in the page, so a police report with a booking photo or a claims file with a vehicle image looks redacted while the picture remains. Redactor rasterises the page so the same detection that runs on images runs on the document.
Can it handle scanned documents with no text layer?
Yes. OCR recognises the text, including on skewed, rotated and low-quality scans, and intelligent character recognition handles handwritten entries.
What about documents that are not in English?
Each text region is classified by script and routed to the engine that reads it, covering Latin including Welsh, Perso-Arabic (Arabic, Urdu, Sindhi, Dari, Pashto), Devanagari, Cyrillic, Chinese and Japanese, and Korean. Routing is per region, so a mixed-script page is handled region by region.
Can it redact a whole column of a spreadsheet?
Yes. Spreadsheets and tables redact by column and by row, so a field of sensitive values is removed in one action rather than cell by cell.
We receive the same form every week. Does it learn the layout?
A reusable template records which regions carry personal data for a recurring document type and applies them to every document of that type, including across a bulk batch. That is what makes batch output consistent between operators.
Which masking styles are available on documents?
Blur, pixelate or a solid fill, chosen per job.
Send Us a Form You Redact Every Week
That is where templates pay for themselves. A bad scan is an even better test.