Redaction, Integrations, Redactor
Human-in-the-Loop Approval Before Anything Leaves the Mailbox
Once software reads a mailbox on an organization's behalf, it is a short step to software that also moves, labels, deletes, replies, forwards and sends the finished production. Human-in-the-loop approval is the usual answer, but the design question underneath it is which of those actions the software may take on its own and which ones wait for a person. The whole release is the subject of our guide to email redaction across a mailbox, and this article draws the line where automation has to stop.
In practice the question reaches two desks at once: counsel or the records officer answers for what leaves the organization, and the security lead has to sign off before an agent touches its mail. Both are usually open to automation, and both want to know where the brakes sit before anything is switched on.
What goes wrong when automation can act on a mailbox
One of the best-known published examples of this risk comes from security guidance for AI applications, and it involves a mailbox. In its entry on excessive agency, OWASP's Top 10 for LLM applications describes an assistant built to summarize incoming email, whose plugin can also send messages. A maliciously crafted email can then trick it into searching the inbox for sensitive information and forwarding it to the attacker.
Every incoming email is external input, so an agent that reads mail is exposed to instructions it never asked for. OWASP's mitigations start with the functions an extension is given rather than with the model, so a summarizer that only needs to read mail should hold no function for deleting or sending it. Its broader advice is to "Utilise human-in-the-loop control to require a human to approve high-impact actions before they are taken."
Header injection is an older version of the same problem, and it matters whenever software composes mail from values that a form or a model supplies. RFC 5322 defines header fields as "lines beginning with a field name, followed by a colon (":"), followed by a field body, and terminated by CRLF." A line break slipped into a subject or a name can therefore end one header and start another, such as a hidden Bcc or a changed Reply-To.
Ordinary human error has not gone away either, and in email it usually takes the form of a send that needed a second look. In 2021 one email from a UK Ministry of Defence team exposed the email addresses of 245 applicants to every recipient, as The Register reported, and the ICO fined the ministry £350,000.
Sorting mailbox actions by what can be undone
Reading mail changes nothing, a label or a move to Deleted Items can be reversed, and a permanent delete or a delivered message cannot be called back. That gives a workable order for deciding who may take each action.
| Action | Can it be undone? | Who should decide |
|---|---|---|
| Reading and searching mail | Nothing in the mailbox changes | An agent, within the mailboxes it was granted |
| Labeling, marking as read, moving to Deleted Items | Yes | An agent only where the workflow's author has chosen to allow it |
| Permanent deletion | No | A person, through a step placed deliberately in the workflow |
| Sending mail | No, once it is delivered | A person, unless an author has deliberately let an agent send |
| Releasing a redacted production | No, once it is received | A person, after review, through an approval step |
When we sorted the mail operations we built for agents on the integration engine that reaches Outlook, sending was the row that took the most thought. We left it off for an agent unless the workflow's author deliberately turned it on. An agent that can send becomes the channel for any instruction hidden in the mail it reads, so that choice has to be made on purpose and by someone who can answer for it.
Permanent deletion is the one mailbox action a user cannot take back, so in any design it belongs with a person. It bypasses trash, and on some mail services it needs a broader grant than reading or labeling does. A mailbox on hold is different, because Exchange Online keeps hard-deleted items in a hidden Recoverable Items folder for as long as the hold lasts, where eDiscovery searches can still reach them. Downloading an attachment and sending a saved draft are also better placed by a person as fixed steps in the workflow, which keeps the model from choosing the target of either.
The same line runs through Microsoft 365 permissions, where reading mail and sending it are separate grants and a pull needs only the first. Our article on Microsoft 365 email redaction explains how an administrator keeps sending and deletion out of the pull.
Where the approval gates belong in an email redaction workflow
Review comes first, straight after detection, where a person confirms, corrects and adds to what the detector found before anything is prepared for release. An approval belongs in front of the release itself, showing exactly what is about to go out, since a reviewer can only approve what the screen shows. What is being approved is often a set of attachments, and our article on redacting email attachments covers what each one should look like by then. A final gate belongs in front of any send, move or delete the workflow makes, because those are the steps that change the mailbox or leave the organization.
An approval that nobody answers needs a defined meaning, or a run can drift toward completion by default. A rejection, or a dialog closed without an answer, should count as a no and stop the run. A person who steps away should come back to a run that is still paused with its state intact, so that walking off neither approves the release nor loses the work.
Whatever tool runs the workflow, the approval decision should be recorded with the release, and that record is what later shows a person looked before anything left. Under Federal Rule of Evidence 502(b), an inadvertent disclosure of privileged material does not operate as a waiver when, among other conditions, the holder "took reasonable steps to prevent disclosure." A recorded approval is part of what a party can point to, and our post on privilege review in eDiscovery covers privilege logs and the Rule 502(d) orders that protect a production in advance.
Guardrails that sit underneath the gates
An approval screen shows only what the software puts on it, so the layer underneath has to stay safe even when a value is hostile. Outbound mail is safest when the server assembles each message from plain To, Cc, Bcc, Subject and Body fields with a proper message builder. Nothing built that way is joined as strings, so a line break inside a value cannot become a new header. A malformed address or a missing recipient should fail before any call is made. Subjects written in accented and non-Latin characters belong in any test set too, since a generic encoder can garble them on the way out.
A connection can drop after a write has reached the mail service but before the answer comes back, which makes retries the other hidden hazard underneath an approval. A client that retries blindly then moves the message a second time, or sends it twice, and the person who approved one send has in effect approved two. The integration engine behind the VIDIZMO workflow retries only operations that are safe to repeat, so a write is sent once and a failure is reported instead of repeated.
Those points become questions for any vendor before an agent is given access to an organization's mail. Ask which mailbox actions the agent can take without a person, and whether the model can choose the target of a write. Ask how outbound mail is built, what a rejection or an unanswered approval does to the run, and whether writes are ever retried after a dropped connection.
Approvals and redaction in the VIDIZMO mailbox workflow
Approval steps belong to the AI Intelligence Hub workflow that reads the mailbox, and redaction stays with VIDIZMO Redactor, as our guide to email redaction sets out. An approval step pauses the workflow and waits for a person before it continues. An agent can also be given a tool that shows an editable preview of what it has prepared and waits for confirmation before going further. A workflow can start when an item is submitted for redaction or when redaction completes, so review and approval follow without anyone watching a queue. Redactor keeps the original by default and releases a redacted copy, so a rejected release leaves the original record exactly as it was.
Legal teams producing email can see how these steps fit a production on the eDiscovery redaction page.
TopicsRedactionIntegrationsRedactor
About the author
Naba Ahtasham is a Product Analyst at VIDIZMO, with three and a half years at the company. An engineer by profession, Naba works on document and email redaction in Redactor, and has also worked on AI Intelligence Hub and AI Live Insight. Much of the job is carrying problems in both directions, from customers, customer success and sales to the engineering team, and back again as a fix that solves what the customer actually needed.
You may also like
Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe
I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...
Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera
A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...
Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video
A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...
See it on your own content
Tell us what you are trying to solve and we will show you how it works on your infrastructure.