Redaction, Audio Redaction, Redactor

Audio PII Redaction: Finding Spoken Card Numbers and Names in a Call

A compliance lead evaluating audio PII redaction tends to ask one question before any other: which parts of a card payment will the tool actually find in a recording? The answer depends on how people say the numbers, where detection runs, and whether the timing of each detection is right to the syllable, since silencing the right words at the wrong moment leaves them audible.

Detection is where call redaction most often succeeds or fails, which is why it has its own place among the failure points in our guide to call recording redaction. The mechanics come first here, because every later choice depends on them.

How a spoken number is found

Nobody finds a card number in a forty-minute call by listening to all of it, so automated detection starts by turning speech into text. The transcript it produces is word-level, which means every word carries its own start time, end time and confidence score. Detection runs on the text, each hit is matched back to the words it covers, and the audio from the first word's start to the last word's end is silenced or beeped. Reading also lets a reviewer check a detection in seconds, since each flagged word leads straight to its moment in the recording.

Where detection runs is a deployment decision with privacy consequences for every recording involved. It can run on AI inside the organization's own deployment, so recordings never leave its infrastructure, or through a cloud service in the organization's own account. The two routes rarely cover the same languages or the same categories of data, so the route has to be settled before a pilot begins.

Why card numbers are the hard case

People say card numbers on the phone in ways that defeat simple pattern matching on a transcript. Every digit left audible is a partial disclosure, so each spoken form below belongs in the set of calls a pilot uses to test detection.

How the number is said Why it is hard What to include in a test set
Sixteen digits in four groups, with pauses A pattern that expects sixteen contiguous digits can miss a number that arrives in pieces Calls where the number is given group by group
"Oh" for zero, "double five" for two fives The spoken words may not match a digit pattern at all Numbers that contain zeros and repeated digits
A correction halfway through a group Both versions of the group sit in the recording A caller who restates part of the number
Expiry date and security code Each is a short, separate item that detection has to find on its own Calls where both follow the card number

Some digits are lost before detection ever sees them, in the steps that tidy a transcript after speech recognition. A common clean-up removes words that fall outside the stretches a speech detector marked as speech, which is how phantom words transcribed from silence are dropped. Speech detectors are weaker on short, quickly spoken numbers, so the same clean-up can remove a real digit, which then never reaches PII detection and stays audible in the "redacted" call. Ask any vendor whether a step between transcription and detection can delete a word containing a number. The answer should be no, even when the speech detector disagrees.

Timing, padding and the edges of a word

Word timings are estimates and speech runs together, so a detection that sits exactly on the word boundaries can still leave the first or last syllable audible. Short detections such as a four-digit PIN or a first name are the most exposed, so their spans are widened slightly at both ends before masking. On a stereo call the widened interval is then masked on both channels, for reasons our article on redacting stereo call recordings explains.

Widening only helps if every output receives it in exactly the same way. A span padded in one file format and not in another means the same call, redacted twice, masks two slightly different stretches, and nobody notices until someone compares them. Redacting one test call into two formats and comparing the lists of redacted segments is a quick way to confirm that the spans match.

Timing also travels through a pipeline as text, and that creates a hazard for anyone building or testing one. Word timings and confidence scores often move between components as decimal text, such as seconds written with a decimal point. On a server whose regional settings use a decimal comma, that text can be read differently, which can shift every redaction away from its word or misread a score. A pilot should therefore run on a server configured exactly like production, regional settings included. The check itself takes minutes: play one second either side of each redaction in a sample and listen for any part of a masked word.

What a confidence threshold does

Every automatic redaction system draws a line between what it masks on its own and what it holds back for a person to review, and the confidence threshold is that line. Detections below it are not masked, and a reviewer should still be able to see them, so the threshold works as a dial between automatic coverage and review workload. A threshold is only as meaningful as the scores behind it, though, and those scores are easy to get wrong in ways nobody notices.

Something as small as a units mismatch between components can distort every score, such as a confidence written as a percentage in one place and read as a fraction in another. A mismatch of that kind can make every match report full confidence, so a threshold meant to hold weak matches back for review sends all of them to automatic redaction instead. The error runs toward over-redaction rather than leakage, but it empties the reviewer's queue of exactly the matches a person should judge. The protection is to test a threshold instead of trusting it, by sweeping it across calls with known PII and confirming that the redaction counts move.

Accuracy figures quoted without your own recordings behind them say little, because line quality, accents and local names differ from one contact center to the next. VIDIZMO works with each organization to calibrate detection to its own calls, and a pilot on real recordings, scored against a set of known PII, measures the result that carries over to production.

What PCI DSS asks of a recording

For payment calls, detection is judged against the PCI DSS rules on sensitive authentication data, which the PCI Security Standards Council applies directly to audio. Its FAQ on audio recordings, quoted in our guide, treats a card verification code kept in any digital audio recording after authorization as a violation. The security code is therefore the item a payment program most needs detection to find.

The council's telephone payments supplement adds a point about the card number itself. It says card numbers must be rendered unreadable anywhere they are stored, including in voice and screen recordings. The primary account number in a recording therefore needs protecting as well, and masking it is a direct way to do that. Detection consequently has to cover the whole payment exchange, from card number and expiry date to security code and PIN, wherever in the call and by whichever speaker each one is said.

Beyond card numbers

Payment data is the most regulated item in a call, but names, street addresses, phone numbers, Social Security numbers and dates of birth turn up in calls that never touch a payment. Each is hard in a different way: names are open-ended, addresses mix numbers with words, and a date of birth looks like any other date until context says otherwise. Callers also spell names and reference numbers out letter by letter for the agent, which spreads one identifier across many short words that detection has to recognize as a single item.

Every organization also has identifiers that no general detector knows about, such as account numbers, case numbers or member IDs. Most tools let these be defined as patterns, whether a regular expression for their structure, a rule that looks for words spoken nearby, or a list of known terms. A pilot should confirm that such patterns are applied to call transcripts as well as to documents, since a pattern that only runs on text files leaves the spoken version untouched.

Language coverage has to be checked per capability, because a system can transcribe many more languages than it can detect PII in. Callers who switch language mid-call, regional dialects, and local street, program or product names that transcription mishears all reduce what detection can find, and each belongs in a pilot's test set. What a misheard word does to the text copies of a call is taken further in our article on PII redaction in call transcripts.

Which payment data Redactor detects, and where

VIDIZMO Redactor transcribes each call to a word-level transcript, detects spoken PII in the text and masks each detection on every channel with silence or a beep. It offers both of the detection routes described above, and they differ in exactly the respects a PCI program cares about.

Detection route Where it runs Payment data detected Languages for spoken PII
The platform's own AI Inside the deployment, wherever it is hosted Card numbers Fewer than transcription covers
Amazon's speech service The organization's own AWS account Card number, verification code, expiry date and PIN English, with Spanish added when a further AWS detection service is enabled

Only the cloud route finds the verification code, so a program that needs the code detected automatically needs that route. One whose recordings cannot leave its own infrastructure has all the more reason to keep the code out of recordings at capture, with keypad entry or pause-and-resume. On either route, short spans can be padded at both ends so word edges are not left audible, and detections below the job's confidence threshold are not masked. Compliance coverage reporting shows how much of a PCI DSS class set, limited to payment and account identifiers, a job actually redacted. Because language coverage differs by route, a pilot should include calls in every language your callers use.

To see the payment-call workflow from detection to release, visit PCI DSS audio redaction.

TopicsRedactionAudio RedactionRedactor

You may also like

Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe

I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...

Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera

A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...

Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video

A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...

See all blogs

See it on your own content

Tell us what you are trying to solve and we will show you how it works on your infrastructure.