Redaction, Document Redaction, Redactor

How to Sanitize a PDF: The Hidden Places PII Survives After Redaction

A litigation team can check every page of a redacted production with copy and paste, find nothing under the bars, and still send the redacted words out in the same delivery. Productions commonly ship an extracted-text file for each document alongside its images so the receiving side can search them, and a text file taken from the native document before redaction carries every word the images withhold. The remedy works for any workflow and any tool: regenerate the text of each redacted document from the redacted version, then search it for the removed words before anything leaves. The rest of the production side, from privilege calls to the privilege log, is in our eDiscovery privilege review article.

The redacted PDF has the same problem on a smaller scale, and our guide to redacting a PDF properly maps where. Earlier saved versions of the file, text written for screen readers, form fields, comments, attachments, thumbnails, properties and bookmarks can all hold words that a redaction applied to the page never touched. To sanitize a PDF is to remove content from those places as well, so that the released copy carries nothing beyond what its redacted pages show. Most of the leaks we have found in real files were sitting in exactly these places.

Earlier saved versions inside the same file

PDF tools often save a change by appending it to the end of the file instead of rewriting the whole document, which is fast and entirely within the rules of the format. The Library of Congress format description notes a feature "introduced to facilitate the incremental updating of a PDF by simply adding to the end of the file." For redaction, the consequence is that the version from before the change can still sit, complete, in the front part of the file.

Appending is not the only way an old version survives, because a tool that rewrites the whole file can still carry objects that no page points to anymore. The image a redaction replaced is a typical example, and it stays unless the tool collects unused objects as it writes. Such objects never display, while a utility that lists every object in the file finds them.

We met exactly that in a redacted PDF whose redaction had been saved as an update appended to the original. Every viewer showed the redacted pages, because viewers follow the newest version recorded at the end of the file. Cutting the file at its first end-of-file marker, which takes a text editor and about a minute, returned the complete unredacted original. A redacted copy therefore has to be written from scratch as a single version, so that nothing from before the redaction goes into it.

A reader can run the same check on any released PDF by opening it in a text editor and searching for %%EOF, the marker that closes each saved version. More than one marker means the file has been added to since it was written, although a file saved for fast web viewing can carry a second marker for innocent reasons. The decisive test is to save a copy that ends just after the earliest marker and open it. If that copy shows pages without their redactions, the history is still in the file.

Accessibility text that speaks the removed words

Tagged PDFs carry text that exists for assistive technology, stored in the document's structure apart from the characters drawn on the page. The W3C's accessibility techniques describe alternate descriptions as "human-readable text that can be vocalized by text-to-speech technology," and show how authors expand abbreviations "by setting expansion text using an /E entry for a structure element." Tagged files can also hold replacement text that tells software which words a run of drawn characters stands for. When a redaction removes the drawn characters and leaves that text in place, a screen reader reads the withheld words aloud, and extraction tools that honor it return them in full.

The file that taught us to look here was a tagged PDF in which one name inside a longer phrase had to be redacted. Removing the drawn characters of the name was straightforward, but the replacement text for the whole phrase still held the complete string with the name inside it, where no visual check would ever show it. Trimming only the redacted words out of that text looks tempting, yet replacement text doesn't map neatly onto the characters it describes. The safer rule is to remove replacement text, alternate descriptions and expansions from every page that carries a redaction, while keeping the pages tagged so assistive technology can still navigate them. A released file can be checked by listening to each redacted page with a screen reader, or by extracting its text with a tool that reads tagged content.

A filled-in form that kept its values three times over

The clearest case we have seen of hidden places adding up was a real, filled-in form sent to us for redaction. A fillable PDF keeps each field's value in the form's own data rather than in the drawing instructions of the page, comments keep their text the same way, and this form also carried accessibility text. A pass that filtered only the page content would have left its values in three places at once, none of them visible to a reviewer.

The approach that holds up is to bring every field value and comment onto the page before anything is removed, so the redaction treats them like any other text. Whatever can't be brought onto the page, such as that accessibility text, gets stripped instead. Handled that way, none of the form's 18 field values survived anywhere in the redacted copy. To check a released form, open its form and comments panes, and treat any field that can still be typed into as a sign that its value may have survived. Form layouts cause detection problems too, from labels printed after their values to numbers that wrap onto a second line, and our article on forms, signatures and stamps works through them.

Attachments, scripts and actions that run on open

Whole files can travel inside a PDF, and in documents that arrive for redaction the attachment most likely to be found is a copy of the original. Nobody redacts inside an attachment while reviewing pages, so a redacted copy that keeps its attachments can hand over the very document it was meant to protect. A PDF portfolio is the extreme case, a file whose cover page may show nothing sensitive while the documents packaged inside it hold everything. A redacted copy should carry no attachments at all, and anything attached that needs releasing should go through redaction as a file of its own.

Document-level scripts, and actions set to run when the file or a page opens, have no place in a released record either, because they can change what a reader sees after the file has passed review. To check a released file, open the attachments pane in a viewer, and use a PDF inspection tool to list any scripts or open actions, expecting none.

Thumbnails, properties and bookmarks

Scanner software often embeds a small thumbnail of each page for quick previews. The thumbnail is made when the page is scanned and never touched by later edits, so it shows the page exactly as it looked before redaction. Viewers build their own previews when a file has none, which makes cached thumbnails safe to remove from every redacted copy, and an inspection tool that lists embedded thumbnails shows whether any remain. Leftover original page images are the other hidden picture a file can keep, and our article on redacting scanned PDFs covers them.

Document properties sit beside the pages, where a title, author, subject or keywords field can repeat a name that was redacted from the text. The Library of Congress notes that the format has "support for annotations, metadata, hypertext links, and bookmarks." It adds that "Version 1.4 and later of PDF can include XMP metadata packages," a second and richer metadata record alongside the basic properties. Clearing the fields in a viewer's properties dialog doesn't always clear that second record, so an old title or author can survive in it after the visible properties look clean. Which properties to remove is a decision each organization has to make, since a record that must prove its chain of custody may need some metadata kept. A written policy naming the properties to mask makes that decision once, instead of leaving it to each reviewer on each release.

Bookmarks, outline entries and link targets can name a section whose text was withheld. That is how a published vaccine supply contract gave away the structure of its redacted clauses, as told in our article on real redaction mistakes. Read every bookmark and hover over every link in the released file before it goes out.

Testing a sanitized copy the way a recipient might

Checks in a viewer show only what the viewer chooses to draw, so the stronger test treats the released file the way a curious or hostile recipient would. Much of the content inside a PDF is compressed, which means a plain text search of the raw file can miss a name that is sitting there in full. Decompressing the file with a PDF utility and then searching the result for every removed word, in the pages, the properties and anything else, closes most of that gap. The rest comes from the way PDFs store text, often as fragments broken up by spacing instructions, so a name can sit in a file split into pieces that a raw search never matches. Searching the text that an extraction tool reassembles covers that case, and the two searches together leave very little unexamined.

We write our own tests on the same principle, as attempts to get the content back out of a redacted file, and a test fails if the attempt succeeds. A vendor asked how it tests its redaction should be able to describe its tests in those terms.

What a Redactor copy leaves out

When VIDIZMO Redactor removes regions from a document, the copy it writes handles each of these places as the table shows.

Place In a Redactor copy
Earlier saved versions and unused objects None, because the copy is written from scratch as a single version holding only the objects its pages use
Accessibility text Replacement text, alternate descriptions and expansions removed from every redacted page, with the pages still tagged
Form values and comments Drawn onto the page before redaction, so they are redacted like other text and the copy is no longer fillable
Attachments Removed, including files pinned to a page as annotations
Scripts and actions Removed, including actions set to run when a page opens or closes
Page thumbnails and leftover images Removed
Document properties Masked where the organization's metadata list names them, with an optional AI check of the rest

Litigation teams can see how these safeguards fit a production workflow on the eDiscovery redaction page.

TopicsRedactionDocument RedactionRedactor

You may also like

Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe

I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...

Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera

A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...

Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video

A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...

See all blogs

See it on your own content

Tell us what you are trying to solve and we will show you how it works on your infrastructure.