Redaction, Audio Redaction, Redactor

PII Redaction in Call Transcripts: Redacting the Text Along with the Audio

A redacted call can play back with every card number silenced while its caption track, sitting beside the audio in the same player, still spells the number out word by word. Nothing about the audio is wrong in that case, so the failure stays invisible to anyone who only listens. PII redaction in call transcripts is the half of the job that tends to be forgotten, even though the transcript is the copy most likely to be searched, exported and kept.

Our guide to call recording redaction counts text copies among the quieter failures, and dealing with them starts with knowing how many a single call leaves behind.

Every copy of the words

A single recorded call can leave its words in more places than the audio file, and each of those places needs an owner who knows it exists.

  • Captions and transcript. The text inside the platform, plus every export of it.
  • Detection results. The record of what was found, which contains the PII text itself.
  • Intermediate transcripts. Files a cloud speech service writes to storage while it works.
  • Temporary files. Anything a redaction job writes to disk along the way.
  • Search indexes. Places where transcript text is indexed, inside or outside the platform.
  • Analytics exports and backups. Copies sent to QA, speech analytics or BI tools, and the backups of everything above.
  • AI-generated text. Summaries, notes and after-call records written from the transcript.

Payment card rules add a reason to redact the transcript, since the PCI SSC's telephone payments supplement says sensitive authentication data that cannot be eliminated must not be queryable. It counts data that a search tool can retrieve as queryable, so on our reading a transcript holding a spoken security code is exactly the queryable copy the council has in mind.

Redacting the text in the same job

Text and audio stay consistent only when both are redacted in the same job, from the same set of detections. Every caption file attached to the call is rewritten alongside the audio, and each PII word is replaced in place by a mask or by the redaction code that explains the withholding. The alternative is to blank the whole caption cue that overlaps a redacted span, which removes the neighboring words as well but leaves no fragment to reassemble.

Numbers and names often straddle two caption lines, because caption cues break on timing rather than on meaning. A card number that begins at the end of one cue and finishes in the next still has to be masked word by word in both. Any keyword removed from the audio has to disappear from the captions as well.

The failure this prevents is easy to picture, because captions are built from the same transcript that detection reads. If the captions sit outside the redaction job, the audio falls silent where the card number was while the caption on screen still shows it. Running both from one set of detections is what keeps them from drifting apart, whichever masking mode a release calls for.

Showing why something was withheld

A mask tells a reader that something was removed, and an exemption code tells them why, which matters when a public-sector contact center releases recordings under records law. A code such as [(b)(6)] can stand in the transcript exactly where the redacted words were. For US federal agencies that placement follows Section 552(b) of Title 5, which asks for the exemption applied to be shown at the place of the deletion, where technically feasible. For internal QA copies a plain mask is usually enough, while codes matter where the text will be read by someone entitled to know the basis for each withholding.

When a caller asks for a copy of their own recordings, word-level masking serves them better than blanking whole cues, since it withholds a third party's details and keeps the caller's own words. The disclosure package for those requests is covered in our article on redacting audio calls for GDPR subject access requests.

What the job leaves behind

Transcription through a cloud speech service leaves a copy outside the redaction workflow unless someone removes it. When transcription runs through a service in the organization's own cloud account, the service writes its transcript of every call to storage in that account. Left alone, an unredacted transcript of every call piles up where nobody reviews it, so ask any vendor whether those intermediate transcripts, and the transcription jobs that produced them, are deleted when processing completes.

A redaction job can also create a sensitive file in the course of its work. To mask many spans in a long call, a job may write a temporary list of every stretch to be masked, which amounts to a plain-text index of exactly where the sensitive speech sits. A job that crashes can leave such a list on disk where nobody would think to look, which is why Redactor deletes it when the job ends, on success and on failure alike.

Detection results and redaction reports raise an access question, because both describe what was said. Ask whether a redacted copy carries forward the detections it masked, and who can open those results and reports. Copies that leave the platform are the organization's to track, since analytics exports, BI extracts, external search indexes and backups hold whatever text they were given. A downstream system receives redacted text only if it takes the redacted item, and only if it takes it after redaction has run. A caption file pulled into a QA tool on the day of the call keeps everything the caller said, however thoroughly the platform's own copy is redacted later.

Summaries and notes written by AI tools are the copy most often overlooked, because they do not look like transcripts at all. A summary generated from an unredacted transcript can restate a card number, a date of birth or an address in words the caller never used, where no search for the original digits will find it. A note pasted from that summary into a CRM record then travels further than the recording ever did. The safe order is to redact first and generate afterwards, so that every summary, note and after-call record is built from the redacted transcript. Anything already generated from an unredacted one should be treated as a copy that needs the same review.

Transcript quality decides audio quality

A surname transcribed as a common noun, or a digit heard as a word, leaves the same gap in the audio as in the text. Detection only ever reads the transcript, and it cannot find a word that was never written down correctly. Transcripts can be corrected by hand before redaction runs, and a program should decide who may correct whose transcripts, since a corrected transcript becomes the basis of the redaction. Our article on finding spoken card numbers and names explains how detection works on that text. What the audio file itself keeps after redaction is the subject of our article on audio redaction and chain of custody.

Checking the text side of a redacted call

The test for all of this takes a known card number and a known name from a test call and looks for them everywhere the text could have gone. Search the redacted item's transcript, then export the captions in every format you use and search those files too, including across line breaks. Open the transcript at each redaction to confirm that a mask or code stands where the words were, and check what the analytics connector or export actually received. After a normal job and a deliberately failed one, look for intermediate transcripts and temporary files as well. A pass means neither the number nor the name turns up anywhere, and nothing that lists or repeats them remains.

The caption file in a Redactor job

A call's caption file is rewritten by VIDIZMO Redactor in the same job that masks its audio. In word mode each flagged word becomes asterisks, or the redaction code in brackets when one is applied. In whole-cue mode every cue that overlaps a redacted segment is masked entirely, and keywords redacted in the job are masked wherever they appear in the captions.

Code lists for US FOIA, UK FOIA and the US Privacy Act come built in, organizations can add their own, and the platform imports and exports eight caption and timed-text formats. Transcripts can be corrected by hand before release, with the right to edit your own content and the right to edit all content granted separately.

To see how a transcript search becomes a masked segment of audio, visit transcript-driven audio redaction.

TopicsRedactionAudio RedactionRedactor

You may also like

Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe

I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...

Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera

A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...

Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video

A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...

See all blogs

See it on your own content

Tell us what you are trying to solve and we will show you how it works on your infrastructure.