Redaction, Document Redaction, Redactor
How to Properly Redact a PDF: What Black Boxes Miss and How to Remove Content for Real
Our first redaction customer at VIDIZMO was a county, in 2022, and most of what it needed redacted was video and audio. That project taught my team that what a reviewer sees or hears is not all a file holds. A muted name can leave enough of its first and last sounds to give it away, and a face that passes behind a pillar can come back into frame unmasked. Documents proved the same thing in their own way, and anyone learning how to properly redact a PDF has to start from that lesson.
A reviewer who pages through a PDF and sees a solid bar over every name has checked what the pages look like, while the recipient gets the whole file. Withheld words can survive under the bar, inside the picture of a scanned page, in an earlier saved version or in form fields and attachments that no page displays. We have found them in every one of those places. The checks below work on any redacted PDF, including one our own software produced.
The document project I keep coming back to is our deployment for a claims adjusting firm that handles claims for insurers across the United States, the United Kingdom and the European Union. A central team of about a hundred people redacts claim files before they go to carriers, brokers, experts and counsel, a few hundred documents a day on a load nobody can forecast. Much of what I say below about real pages comes from those files.
What a black box misses, and how to properly redact a PDF
A black box drawn with the shape or highlight tool of a PDF editor is an annotation on top of the page, and the text underneath stays where search, copy and extraction still reach it. Built-in redaction features usually mark an area first and apply the redaction second, and a file saved between the two steps looks finished while still holding everything. A 2023 study by researchers at the University of Illinois found "6,541 nonexcised redacted names in 710 US court documents," meaning names whose text was still sitting under the bar.
The UK National Archives sets the standard in its redaction toolkit: "The information must be completely removed from the bit stream, not simply from the displayable record." Properly redacting a PDF means nobody can recover the withheld words from the released file by any means, whatever its pages look like. How the mark looks still matters for text, which should get a solid fill, since a 2016 study concluded that "mosaicing and blurring, despite their widespread usage, are not viable approaches for text redaction."
Can a redacted PDF be unredacted?
A redacted PDF can be unredacted whenever the redaction covered content instead of removing it, usually by copying the text from under the bar, deleting the annotation, extracting the page image or opening an earlier saved version. Content removed from every place a PDF keeps it can't be brought back, though the layout around the gap can still hint at it. The Illinois researchers estimated that a typical PDF authored in Microsoft Word "leaks about 13 bits of information about a redacted surname" through the positions of the surrounding characters. By their account, that is enough to identify one individual from 8,000 candidates. The risk is highest where the pool of possible names is small, as with the staff of one office, so those releases need a reviewer who asks what each page's layout gives away.
Where a PDF keeps content besides the page
A PDF is a container, and the page on screen is assembled at display time from text, images and drawing instructions stored separately, alongside data that no page shows. Inside, the file is a graph of numbered objects, such as pages, fonts, images and the streams that draw them, with an index at the end saying where each one sits.
Editing a PDF rarely takes anything out of that graph, because most editors save a change by writing a new version of the object and pointing the index at it. Many append that new version to the end of the file, so the original object stays where it was. Even a tool that rewrites the whole file can carry objects that nothing uses anymore, unless it collects them as it writes. A viewer draws only what the pages reference, so the original is invisible on screen, while anyone who reads the file's objects directly can get it back. That's why I treat redaction as a deep clean of the file rather than an edit to its pages. The table lists the places we check and why a page-level redaction misses each.
| Where content survives | Why a page-level redaction misses it |
|---|---|
| Text under a drawn box | The characters stay in the page content, where search and copy find them |
| Scanned pages | The box covers the picture, while the page image, its layers and the hidden recognized text still hold the words |
| Earlier saved versions | An update appended to the file leaves the pre-redaction version inside it |
| Objects nothing points to | An edit can leave the original text or image in the file, never drawn but readable by any tool that lists the file's objects |
| Accessibility text | Text written for screen readers can spell out the removed words |
| Form values, comments and attachments | They are stored apart from the page, and an attachment is often the original itself |
| Document properties and XMP metadata | Author, title, subject and keyword fields, and a second XMP record beside them, can repeat a name redacted from the text |
| Thumbnails and bookmarks | A scan-time preview or an outline entry can show or name what was withheld |
One redacted PDF we examined still carried its unredacted original as an earlier saved version, and one filled-in form kept its values in three places the page never showed. Each hiding place gets its own section in how to sanitize a PDF, with the file that taught us to look there and a check any reviewer can run. The rule we took from those files is to write every redacted copy from scratch as a single new file, carrying only the objects its pages still use, so nothing an editor left behind travels with it. Form values are drawn onto the page before anything is removed, and attachments, scripts and thumbnails are left out. Document properties such as author, company and dates are checked against the organization's metadata list, with AI detection as a second pass, and flagged values are masked. A property nobody listed and detection doesn't flag stays as it was, so I'd review that list as carefully as the redaction policy itself.
Scans, photographs and the pages detection misses
On a scanned page the words are part of a photograph of the paper, so a box drawn on top only adds a layer, and extracting the page image gets them back. Searchable scans also hold the words as invisible recognized text, and compressed scans often split each page into stacked images. On one layered scan we examined, a painted box had left every layer beneath it readable. A redaction that holds blanks the pixels under each region in every image and removes the hidden text in the same pass. The steps and checks for that are laid out for anyone redacting a scanned PDF.
Removal only protects what detection found, and the claims firm's files show how real pages defeat detection built for clean, upright, single-language print. A European claim can carry a police report in the local language, a policy schedule in English and a medical note in a third. Detection usually runs in one language at a time, so names in the other two can slip past it. Multilingual and right-to-left documents fail even earlier when the OCR wasn't built for their script, since a reader trained on Latin letters gets nothing usable from an Arabic or Urdu page.
Claimants photograph paperwork at whatever angle they held the phone, and on a sideways or skewed page a box mapped to the wrong frame lands on the wrong words. Rotation has a dependable fix while skew and handwriting still need a person, as what OCR misses on handwritten, rotated and skewed scans explains. Photographs pasted into claim forms, including pictures of identity documents, carry faces and license plates that text search can't see, so they have to be found as objects on the page. Forms and signature blocks break detection too, from labels printed after values to notary seals that read as graphics. Redacting signatures, forms and stamps in a PDF starts from the policy question of which signatures need withholding at all.
At the firm, AI detection and the firm's own rules run in one pass, because neither covers the other's ground. Rules catch the identifiers known in advance, such as claim reference formats and policy number patterns, while detection finds the names nobody listed and the faces in the photographs. The team still makes the calls that no rule or model should make alone, such as a handwritten note in a margin, or a carrier that wants one category of information kept when another wants it removed. We calibrate detection to each organization's own material and measure it in a pilot, and a release plan should still budget reviewer time for the pages that need judgment.
Word files and spreadsheets come back as PDFs
Today every Office document that goes through VIDIZMO Redactor, whether Word, Excel or PowerPoint, comes back as a redacted PDF, and I count that as a limitation. The conversion does buy something, because a native Office file keeps content its pages never show, such as tracked changes, comments, hidden worksheets and speaker notes. A PDF copy fixes what every page shows and gives documents, spreadsheets, slide decks and saved emails one removal path. The National Archives toolkit also recommends PDF as a format for redacted copies, warning that "Some binary formats may allow changes to be rolled back."
Customers have told us that spreadsheets are where this hurts most, since whoever receives a workbook usually wants to sort, filter and total it, and a PDF hands them pages instead. The practical answer is to agree the format of the redacted copies before redaction starts. On a records request that means agreeing it with the requester, and on a production it belongs in the discovery protocol, so any disagreement surfaces before a single file is processed. What a conversion can leave out, and how a litigation production differs from a records release, is weighed in how to redact a Word document or spreadsheet.
My team made the opposite call for DICOM medical images, because the applications that receive a study expect DICOM. A study redacted in Redactor stays a DICOM file, with the text burned into its pixels masked and its header de-identified in the same pass.
Marking what was withheld, and keeping the record
A public records office usually redacts against a statutory deadline, with each request logged in a case-management system. The federal Freedom of Information Act gives a US agency 20 days, not counting weekends and legal public holidays, to decide whether to comply. A UK public authority must answer "not later than the twentieth working day following the date of receipt" under section 10 of the Freedom of Information Act 2000.
Both laws also expect the reason for each withholding to be visible to the person who receives the release. The US Act requires that, where technically feasible, "the amount of the information deleted, and the exemption under which the deletion is made, shall be indicated at the place in the record where such deletion is made." It excuses the marking only where the indication itself "would harm an interest protected by the exemption". When a UK authority withholds information, its refusal notice must specify "the exemption in question" under section 17. A code printed on the redaction travels with every copy of the record, where a separate log can be lost, and a fixed list of codes keeps analysts consistent.
When a requester sues, a US agency commonly defends its withholdings with a Vaughn index. The Justice Department's guide to the FOIA explains that the decision behind it requires agencies "to correlate each withheld document (or portion thereof) with a specific FOIA exemption." In Redactor each redaction carries its code on the page, from the default US FOIA, UK FOIA or US Privacy Act lists or from an organization's own list. The redaction report is a redaction log rather than a Vaughn index. It records what was redacted, by whom and when, and exports as CSV, so counsel drafting the index works from a record built while the redaction was done.
Before release: labels, originals and the questions to ask
A redaction job can fail quietly in a way no visual check catches, by handing back a file that nothing was removed from under a label saying it was redacted. That label is what stops the next person from checking, so we built Redactor to skip a document job that masks none of its regions, with a stated reason and no redacted copy or label. A job that can't read its detection results fails outright instead of reporting nothing to redact, and one that fails partway leaves the original as it was.
The original has to survive too, because a withholding can be challenged and someone must be able to compare the release with its source. The National Archives toolkit says to "Never redact the original or master version of an electronic record," and Redactor keeps the original and writes each redaction as a separate copy unless an organization sets a different policy.
A few process checks catch what no single-file test can, starting with opening the released file in a fresh session on another machine and comparing its page count with the original's. Confirm that the file going out is the redacted copy rather than a similarly named original, and sample the difficult pages for boxes on the wrong words. Put a second reviewer on any release where one missed name would do serious harm.
What to ask of a redaction tool
The questions I'd put to any vendor, ours included, follow from all of this, and the answers show quickly whether a tool removes content or only covers it:
- Does it delete the text and blank the pixels under each region, in every image and layer beneath it?
- Is the redacted copy written as a single new file, and what happens to screen-reader text, form values, attachments, scripts and document properties?
- Which scripts can its OCR read, which languages does its detection cover, and how does it handle rotated pages?
- What does it hand back when a job removes nothing or fails partway, and does the original survive?
- In what format does a redacted spreadsheet come back?
- Can files arrive from our request or review system and go back with their redaction record?
What to license, and where it runs
For the document work in this guide you license Redactor, which runs on the VIDIZMO platform that holds your files, users and access rules. The same product redacts video, audio, images and DICOM studies. It can be hosted by VIDIZMO, run in your own cloud subscription or on your own servers, or run air-gapped with every AI model running locally.
The eDiscovery review platforms litigation teams work in connect through AI Intelligence Hub, a second product licensed separately, which runs those connections and the workflows around them. A records or case management system needs no second product to work with Redactor, since it can hand files over through the platform's REST API and hear back through a webhook when a job completes. Redacted copies can be exported to AWS storage or SharePoint.
Formats, detection methods and output options are listed on the document redaction software page.
TopicsRedactionDocument RedactionRedactor
About the author
Farooq Khan is the co-founder and CTO of VIDIZMO, where he leads the engineering, product, and AI strategy behind its platform for making sense of unstructured media. He builds applied and generative AI that turns organizations' video, audio, images, and documents into searchable, governed, and usable intelligence at enterprise scale. Over the past 20 years, he has built systems that capture, scale, and now understand media, from voice logging platforms to large scale commerce to VIDIZMO's AI platform. Today that platform is trusted by global enterprises and government agencies alike.
You may also like
Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe
I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...
Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera
A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...
Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video
A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...
See it on your own content
Tell us what you are trying to solve and we will show you how it works on your infrastructure.