Home / Document redaction software

Document Redaction Software

Three ways to find what to remove — keyword, pattern, or AI detection — across PDF, Word, Excel, PowerPoint, Visio, OneNote, CAD, email and HTML. Including the photograph on page four that text-only tools cannot see.

(b)(6) EMBEDDED PHOTO (b)(7)(C) Text and the photograph inside it, both redacted

Why Teams Move

4M pages

Redacted a year on a single deployment

6

Script families read, from Latin to Perso-Arabic

63

Document formats in, no conversion step

What Makes Documents Hard

Scans are rarely clean

Photocopies arrive skewed, rotated, faint, or photographed on a phone at an angle. Text with no text layer behind it. OCR has to cope with the document you were sent, not the one you wish you had.

Documents are not just text

A police report carries a booking photo. A claims file carries a vehicle image with a plate in it. A benefits application carries a scanned ID. A text-only tool redacts the character runs and leaves the photograph — and the redaction looks complete.

One page, several types

Printed text, handwriting, a signature, an embedded image and a table, all on the same sheet, each needing different handling.

One column, 400 rows

Spreadsheets need redaction by column and by row, not cell by cell.

Not every document is English

Each text region is classified by script and routed to the engine that reads it — Latin including Welsh, Perso-Arabic covering Arabic, Urdu, Sindhi, Dari and Pashto, plus Devanagari, Cyrillic, Chinese and Japanese, and Korean. Routing is per region, so a mixed-script page is handled region by region.

The same form, every week

It arrives on a schedule and gets marked up by hand every time.

Formats nobody planned for

, and metadata that survives the redaction unless it is stripped on export.

How to Find What to Redact

Keyword

Specific words or phrases, wherever they appear

Pattern

Regular expressions and predefined templates such as SSN or credit card

AI-assisted

PII, text recognised through OCR, and objects such as faces and signatures

Templates for recurring types

Where documents share a layout, the regions carrying personal data sit in the same places. A template records those regions and the classes to detect, then applies to every document of that type, including across a bulk batch. That is what makes a queued batch consistent rather than dependent on which operator ran it.

What You Can Redact

Printed text

PII, PHI, CCI and financial data in the document body

Scanned text

Pages with no text layer, recognised through OCR

Handwriting

Handwritten entries, through intelligent character recognition

Signatures

Detected as a class of their own

Embedded images

Faces, plates and vehicles inside photographs on the page

Table columns and rows

A whole field of values in one action

Non-Latin scripts

Perso-Arabic, Devanagari, Cyrillic, CJK and Korean

Keywords and patterns

Specific phrases, or regex and templates such as SSN

Document metadata

Attributes carried in the file, stripped on export

Proof

VIDIZMO

86.1%

Idox plc

32.8%

ReadyRedact

30.7%

The Department for Work and Pensions in the UK runs the largest documented deployment: roughly 5,000 users processing 4 million pages a year on a single installation. That answers whether the architecture holds at document volume rather than in a demo.

It also produced the most useful competitive evidence available, because the scoring above was done by the buyer against public-sector requirements — not by a vendor writing its own comparison table.

Compared With Adobe Acrobat

Acrobat Redactor
One file at a time, one person at a time Bulk batches, queued and worked overnight
Text only; an embedded photo is opaque to it Page rasterised, so faces and plates inside PDFs are detected
Mark up each recurring form by hand, every time Templates apply the pattern to every document of that type
Cell by cell on a spreadsheet By column and by row
Needs a text layer before it can help OCR and handwriting recognition built in
No exemption codes, no coverage report Statutory code on every redaction, printed on the output

Where It Runs

Shared SaaS, dedicated cloud, your own private cloud, on-premises, hybrid, or fully air-gapped.

FAQ

Document Redaction Software questions, answered

Is this different from using Adobe Acrobat?

Acrobat redacts one file at a time and works on the text layer. It cannot see a photograph embedded in the page, so a police report with a booking photo or a claims file with a vehicle image looks redacted while the picture remains. Redactor rasterises the page so the same detection that runs on images runs on the document.

Can it handle scanned documents with no text layer?

Yes. OCR recognises the text, including on skewed, rotated and low-quality scans, and intelligent character recognition handles handwritten entries.

What about documents that are not in English?

Each text region is classified by script and routed to the engine that reads it, covering Latin including Welsh, Perso-Arabic (Arabic, Urdu, Sindhi, Dari, Pashto), Devanagari, Cyrillic, Chinese and Japanese, and Korean. Routing is per region, so a mixed-script page is handled region by region.

Can it redact a whole column of a spreadsheet?

Yes. Spreadsheets and tables redact by column and by row, so a field of sensitive values is removed in one action rather than cell by cell.

We receive the same form every week. Does it learn the layout?

A reusable template records which regions carry personal data for a recurring document type and applies them to every document of that type, including across a bulk batch. That is what makes batch output consistent between operators.

Which masking styles are available on documents?

Blur, pixelate or a solid fill, chosen per job.

Send Us a Form You Redact Every Week

That is where templates pay for themselves. A bad scan is an even better test.