Redaction, Integrations, Redactor

PII in Emails: Bodies, Signatures and Quoted Replies

Ask where the personal information in an email sits and most people point to the body, where someone typed an account number, a diagnosis or a home address into a reply. In a real thread, PII in emails is spread much wider than that. The header block names everyone on the conversation, and addresses built as first.last@ spell out a person's name before anyone reads a word of the message.

The signature adds titles, direct lines, mobile numbers and office addresses, sometimes with a photo or a scanned signature. Every reply carries a quoted copy of what came before, so one sentence can appear six times in six slightly different layouts. A forwarded block brings whole earlier messages along, each with its own header lines. Our guide to email redaction across a mailbox follows a request from the pull to the release, while the map below stays inside a single message and its thread.

Where PII sits in one message

Personal information turns up in more places in a message than a reader's eye goes, and the places a reviewer skims are the places it survives. The table maps each location, what it typically holds and why it gets missed.

Where in the message What it typically holds Why it gets missed
Header block Display names and addresses of the sender and every recipient Reviewers read the body and skim the lines above it
Subject line Names, case numbers and the reason for writing It repeats in every reply and rarely gets read as content
Body Account numbers, diagnoses, addresses, pasted tables and screenshots Text inside a screenshot is a picture, so a text search never sees it
Signature block Name, title, direct line, mobile number, office address, photo, scanned signature After the tenth message it reads as boilerplate
Quoted replies Every earlier message, plus the "On [date], [name] wrote:" line Each copy is reflowed, so a search built on one layout misses the others
Forwarded blocks Whole earlier messages with their own From, To and date lines They look like body text but carry a second header block
Routing headers, in the original file only Host names and IP addresses along the path the message took Nobody reading the thread ever sees them

When the original message file is what gets released, every line of its routing headers travels with it unread. Microsoft's message reference describes the header set as including "message headers indicating the network path taken by a message from the sender to the recipient."

Asking only for the fields a step needs, such as sender, recipients, subject and date, keeps those lines out of any tool or agent that reads the mail. A release follows the same logic when it renders only what a reader of the message could see, instead of shipping the original file. What a mail service returns for a message, and how to scope a pull, is covered in our article on Microsoft 365 email redaction. Attachments carry personal information of their own, which our article on redacting email attachments takes format by format.

What counts as personal data, and what can stay

UK data protection law counts the contact details in a work email as personal data, even when the message itself is about business. In the words of the ICO's guidance on what personal data is, "A name and a corporate email address clearly relates to a particular individual and is therefore personal data."

Whether those details are then withheld is a policy decision, and the answer differs between a FOIA release and a subject access response. Our post on GDPR redaction for subject access requests covers the subject access process. Whatever the policy, a release only holds together if it is applied the same way in every message of the set. Email chains mix the requester's own data with colleagues' details and other people's, and every withholding has to be defensible to a regulator or a court. A name withheld in one reply and released in the next tells the requester exactly what the redaction was hiding.

Quoted replies are where consistency breaks

Each reply in a thread repeats earlier text in a new layout, with lines rewrapped, prefixed with markers or indented, and sometimes converted from HTML to plain text along the way. A name redacted in one message can then survive in the quoted copy two messages later. Microsoft Graph even exposes the difference, describing a property called uniqueBody as "The part of the body of the message that is unique to the current message." The copy that gets released contains both parts, so the redaction has to land on every copy.

Over-redacting quoted history is an error too, and in a public records release it can matter as much as a missed name. Under FOIA a whole string of emails can be a single record, as our guide to email redaction explains, and non-exempt text inside a responsive record cannot be removed as non-responsive. Quoted material that no exemption covers therefore stays in, even when it repeats something the requester already has.

Because detection reads the whole rendered page, each quoted copy is its own occurrence at its own position instead of a duplicate to be skipped. A vocabulary list of the names that must go catches every copy of them, however the thread has reflowed the text. A reviewer then compares the copies of each redacted passage side by side before anything is released.

Signatures and the details that repeat in every message

A direct line or a mobile number appears in every message its owner sends, which makes the signature block the densest and most repetitive source of contact details in a thread. Those numbers are best found with patterns that look at the words around them, such as "Mobile", "Direct" or "Cell", because a bare ten-digit number could be almost anything. The label can come after the number as well as before it, as in a line reading "555 0100 (mobile)", so a pattern has to look both ways.

The organization's own name, main number and street address appear in every signature, and when they are flagged every time, they bury the finds that matter under hundreds of identical hits. They belong on an exclusion list, and two properties of that list decide whether it works. An entry should match every capitalization, since the same name appears in title case in one signature and in capitals in a letterhead. An exclusion should also win over any detection that would otherwise match the term.

A scanned handwritten signature pasted into a body or a signature block is an image rather than text, so no text search will ever find it. It has to be found as an object on the page and masked like any other region.

Patterns that slip past detection, and the checks that catch them

A name that appears only inside an email address, such as jdoe@, never appears as a name at all, and it is one of several email patterns that defeat detection whatever tool does the work. An address or a phone number split across a line break in quoted text can read as two fragments, neither of which looks like the whole. A thread that switches language partway through can defeat a detector configured for one language, even when every name in it is spelled correctly. Each of these is easy to reproduce in a test thread before a tool is trusted with a release, and our PII redaction software guide covers how detection works more generally.

The release is what a requester will actually read, so the searches that catch these misses run on the released text instead of the redaction tool's view of it. Each of the searches below takes only minutes on a sample:

  • Search the released text for "@" and for the organization's own domain, including the domain broken across two lines.
  • Search for the part of each redacted address before the "@" on its own, such as jdoe.
  • Search for each redacted name, then compare every quoted copy of the passage it appeared in.
  • Search for "wrote:" and "From:" to find every quoted and forwarded block, and read their header lines separately.
  • Read every subject line in the set on its own, apart from the message bodies.
  • Zoom into each screenshot in a message body and confirm that the text inside it was reviewed.

What VIDIZMO detection looks for in a message

In VIDIZMO Redactor, the header block of an .eml or .msg file is rendered with its body, so the names and addresses shown there go through the same detection as the rest of the page. Detection covers classes that include person, email address, phone number, URL, IP address, username, organization and location, alongside country-specific identifiers. It looks for a field label after a value as well as before it, and a pattern that breaks across a line is matched in both parts, each of which is masked.

Organization-specific identifiers are added as regex, context-word or vocabulary patterns, and handwritten signatures are detected as objects on the rendered page and masked. A reviewer confirms, corrects and adds to every detection before anything is released.

Teams releasing correspondence can see how Redactor handles whole email threads, from headers and signatures to quoted replies and attachments.

TopicsRedactionIntegrationsRedactor

You may also like

Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe

I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...

Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera

A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...

Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video

A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...

See all blogs

See it on your own content

Tell us what you are trying to solve and we will show you how it works on your infrastructure.