Redaction, Document Redaction, Redactor

How to Redact a Scanned PDF: Removing the Pixels Instead of Covering Them

Extract the page images from a scanned PDF that someone redacted by drawing black boxes, and the names under the boxes come back while the page on screen still looks perfect. The extraction takes about a minute with an image-extraction tool and needs no skill beyond knowing that such tools exist. On a scan the words are pixels in a photograph of the paper. To redact a scanned PDF, the removal has to reach every copy of those pixels and every copy of the recognized text, including copies that no viewer displays. Our guide to properly redacting a PDF maps every place a PDF keeps content, and the image side of that map is where scanned files go wrong.

Finding the words on a scan is a separate job from removing them, and it has to happen first. OCR makes the text of a scan findable, and our article on handwritten, rotated and skewed scans covers the pages OCR reads badly. The procedure below covers the removal, which is what decides whether a redaction on a scan holds.

How to redact a scanned PDF, step by step

The steps work with any redaction tool that removes content instead of drawing over it, and the checks in steps five to eight work on any redacted scan, whichever tool produced it.

  1. Keep the scan as it arrived, and do all the work on a copy.
  2. Run OCR on the copy if the scan has no text layer, so detection and search can find the words, and look over pages carrying handwriting or stamps by eye.
  3. Mark every region to withhold, then apply the redaction with a tool that deletes the recognized text and blanks every pixel beneath each region in one pass.
  4. Save the result as a new file instead of an update to the old one, so no earlier version of a page survives inside it.
  5. Extract every image from the redacted file and look at each redacted area, checking every layer of a layered page.
  6. Search the text layer for each removed word, and for a likely misreading of it, since OCR often turns a name into something close to it.
  7. Compare the number of images in the file before and after redaction, and on pages that share a letterhead or logo, confirm that the redaction appears only where it was meant to.
  8. Render each redacted page next to its original at the same size, and check that the only differences sit inside the redaction boxes and that rotated pages kept their orientation.

The sections below explain the traps behind steps three to eight, each of which we have met in real files sent for redaction.

Two copies of every word on a searchable scan

A scan is stored as a picture, and the W3C's accessibility guidance notes that "A document that consists of scanned images of text is inherently inaccessible because the content of the document is images, not searchable text." Searchable scans answer that by running OCR, which in the W3C's words "converts images of words and characters to actual text," and storing the result as an invisible layer behind the picture. The page then carries every word twice, once as pixels and once as hidden text that search, copy and screen readers use.

Most digital workflows scan and OCR paper long before anyone decides what to withhold, so nearly every scan that reaches redaction carries both copies. Blanking the pixels while leaving the hidden text leaks as much as leaving both, and that is why step three asks for a single pass. Where the hidden layer came from older OCR software it may not line up with the picture, and a region drawn over the visible words can then miss their hidden copy. A quick way to spot a misaligned layer is to select a line of text in a viewer and watch where the highlight falls. If it sits above, below or beside the printed words, the hidden text and the picture disagree about where each word is. The safer course then is to run OCR again on the copy before any region is drawn.

Paper blacked out with a marker before scanning has a different weakness, covered in our article on redacting pens.

Layered scans store one page as several images

Many compressed scans store a single page as a stack of images that the viewer combines into one ordinary page. A typical stack holds a background, a sharp black-and-white mask that carries the text, and a foreground that gives the text its color. The technique is Mixed Raster Content, and the ITU recommendation that standardizes it describes "segmentation of the image into multiple layers (planes)," noting that "Recombining the layers in a prescribed manner regenerates the original image." The point of the split is compression, since the text mask can stay sharp while the background and foreground are compressed heavily, and many archives digitized at volume use it.

We have seen what a layered page does to a redaction aimed only at the picture on screen. On one layered scan, a box painted over the page left every layer beneath it untouched, and extracting the layers returned the words in seconds while the page looked flawless on screen. VIDIZMO Redactor blanks the region in every image beneath it, background and text mask alike, so no layer can reassemble what the page no longer shows. On a layered test scan, both layers came out blank under each region.

Leftover page images

Blanking pixels means writing a new image for every redacted page, and what happens to the old image decides whether the redaction holds. We have examined a redacted 40-page scan that still carried all 40 original page images, because the save that wrote the new ones never cleared out the old. Nothing referred to the originals anymore, so no viewer displayed them, yet any tool that walks through the objects in a PDF could pull them out. An inspection tool that lists a file's image objects shows the originals if they are still there, which is what the count in step seven is for. Redactor writes a redacted file as a single new version that contains nothing its pages no longer show.

Scanner software can leave one more picture of each page behind, a small thumbnail embedded at scan time, and our article on sanitizing a PDF covers it along with the other hidden places.

Faces, photographs and ID cards on a scanned page

Scanned case files often carry more than text, such as a photocopied driver's license in a benefits application, a printed photograph clipped to an incident report, or an identity card copied onto a form. A face on those pages is pixels in the same picture as the words around it. It needs the same treatment as a name, found on the rendered page and then blanked in every image beneath the region instead of covered. Detection of faces on a scan proposes regions for a reviewer to confirm, because a photocopy of a photocopy can blur a face past the point where software recognizes it.

Redactor's detection runs on the rendered page, so it finds faces and signatures on a scan the same way it finds them in a born-digital document. Signatures raise a policy question as well as a technical one, which our article on signatures, forms and stamps takes up.

Pages that share one image

A letterhead, a logo band or a separator sheet that repeats across pages is usually stored once, with every page that shows it pointing at the same copy. A redaction on one of those pages can go wrong in either direction. Editing the shared image in place punches the same hole into every page that uses it, destroying content nobody meant to withhold. Leaving it alone has the opposite effect and keeps the original visible on the page that needed the redaction.

The safe behavior is to give the redacted page its own copy of the shared image with the region blanked, while every other page keeps pointing at the untouched original. The check in step seven shows whether a tool does that. When the repeated image itself carries something that must be withheld, such as a staff member's name in a letterhead, every page that shows it needs a redaction of its own.

Scans from print and publishing workflows

PDFs that come out of print production often use image color formats built for presses, and spot-color and palette images can't always be edited pixel by pixel. A redaction pass that expects ordinary images then either stops with an error or can't blank the region at all. We worked through this with an eight-page newspaper issue from a print workflow that carried 134 images, many of them in those formats. Dropping each affected image would have removed a photograph along with the small area that needed redacting. Flattening the page into a picture would have destroyed its text layer, about 11,000 characters on one page. Redactor converts only the images a redaction actually touches, and leaves the text layer and every untouched image as they were.

Flattening a page into a picture also runs into court rules on electronic filings, such as the California appellate rule that requires "text-searchable portable document format (PDF)" unless conversion cannot practicably be done. The flatten-to-image approach that some archives recommend therefore suits a small release better than a filing or a production that has to stay searchable.

Where Redactor fits on a scanned page

On a scanned page, Redactor removes each region from the text layer and from every image beneath it in the same pass. The region can come from OCR, from AI detection or from a box an analyst drew, and nothing outside it changes, including on pages stored sideways.

Records teams preparing scanned material for release can see the wider workflow on the public records redaction page.

TopicsRedactionDocument RedactionRedactor

You may also like

Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe

I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...

Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera

A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...

Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video

A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...

See all blogs

See it on your own content

Tell us what you are trying to solve and we will show you how it works on your infrastructure.