Redaction, Audio Redaction, Redactor
Call Recording Redaction at Scale: What Actually Breaks and How to Handle It
I led our call recording redaction deployment at a county social services contact center, where about 7,000 calls come in on an average working day, and as many as 10,000 on a busy one, and each has to be redacted the same day. The calls are short, recorded in stereo and delivered with a metadata file each. Callers give their names, dates of birth, addresses and card details, because that is how the agency identifies and serves them. Before the pipeline could take on that daily flow, it also had to clear about two million recordings stored before the program began, none of them ever analyzed because none could be shared.
Most people responsible for a contact center's recordings have already decided to redact them and are now working out how. The demonstrations they see usually run on a handful of clean sample calls, while the failures that matter turn up in production. They surface at volume, on the second channel of a stereo file, at the edges of a masked word and in the text that travels with every call. What follows is what we ran into at the county and what we decided, along with what I'd ask any vendor to prove on your own recordings.
What makes call recording redaction hard
Call recordings are short and numerous, and the personal data in them is spoken the way people actually talk on the phone. Card numbers arrive in groups, with pauses and corrections along the way. At the county, callers regularly spelled their names out letter by letter and read house numbers digit by digit for the agent, which spreads one identifier across many short words. Many called from a roadside or a bus, so background noise sat under exactly the words that mattered.
A call rarely travels alone, and the files that come with it describe the same conversation in other forms. The audio usually arrives with a metadata file from the contact center platform. Once the call is transcribed, a transcript and a caption track repeat the conversation in text that anyone can search.
Why an organization redacts its calls in the first place shapes what a good redacted result has to look like. Payment rules demand that card codes don't survive in storage, and a records release demands that every withholding can be explained to the requester. Sharing recordings with an outside analytics or QA provider demands that nothing identifying leaves at all. That was the county's reason: it wanted its calls analyzed for service quality, the recordings couldn't reach the analyst as they stood, and redaction became the pipeline that let them go.
PCI DSS rules for recorded payment calls
For payment calls the rules are specific, because the PCI Security Standards Council has answered the recording question directly in FAQ 1210, last updated in June 2025. Requirement 3.3.1 prohibits storing sensitive authentication data after authorization, even when it's encrypted. A card verification code kept in any digital audio recording after authorization, a .wav or .mp3 file included, is therefore a violation. The FAQ says technology that suppresses or redacts audio during data entry should be enabled wherever it exists. Where the code can't be kept out of the recording, it should be securely deleted immediately upon authorization.
Keypad entry and pause-and-resume keep card data out of new recordings, and they're the right first control, but they leave the archive already on disk exactly as it was. The council's telephone payments supplement accepts that a properly implemented pause-and-resume solution can take the recording and storage systems out of scope. It also says the technology doesn't reduce PCI DSS applicability to the agent, the agent desktop environment or any other systems in the telephone environment. Because manual pausing depends on agents remembering, the supplement encourages organizations to ask their call center operator how card codes are removed from recordings, preferably automatically. Where pause-and-resume is in use, it also recommends checking recordings for card data regularly, preferably weekly.
For a contact center, those rules turn into a handful of decisions that belong in writing rather than in a default setting. Card codes should be stopped at capture wherever the platform allows it, and the codes already sitting in the archive have to be found and redacted. An unredacted original that caught a security code after authorization shouldn't survive unless a documented legal or evidential requirement holds it. Without one, the redacted copy should replace the original, or the original should go to a recycle bin with a short retention period.
Which detection route you choose decides which payment data is found automatically, and the difference matters to a PCI program. In VIDIZMO Redactor, detection on the platform's own AI inside the deployment finds card numbers. The verification code, expiry date and PIN need the route through Amazon's speech service in your own AWS account. A contact center whose recordings can't take that route has all the more reason to keep codes out of recordings at capture.
Inside the customer's account, from ingest to export
The first requirement the county set had nothing to do with detection, because its unredacted recordings were not allowed to leave its own AWS account. We deployed Redactor into that account, so our software ran on the county's own infrastructure and data, and no unredacted call crossed the boundary. A hosted service that needs calls uploaded to it fails a requirement like that before its accuracy is ever tested, so settle where your recordings may go before you compare tools.
Getting calls in was the next problem, since a contact center platform typically writes a call's audio and its metadata as separate files that land in storage at different moments. A pipeline that takes whatever is present creates calls without metadata and metadata with nowhere to go. A pickup run cut off partway also has to resume without taking any call twice, and a clean-up step must never delete a source file before its item exists. Redactor takes a call in only once all of its files have arrived. How a call's audio and metadata can be ingested together from S3 is worked through on Amazon Connect, because AWS documents that platform's storage behavior in unusual detail.
Sending calls back out mattered as much, because the analysis on the far side expected each redacted call in the same folder structure and WAV format, with its metadata file beside it. When Redactor redacts detected PII in an uncompressed PCM WAV call, it overwrites only the bytes inside each redacted span and copies every other byte unchanged, header included. The downstream system receives exactly the format it accepted before, and anyone can compare the two files to show that nothing outside the spans moved. The case for changing only the bytes that carry the PII also explains why compressed recordings can't be handled that way and how to verify any tool's output.
How we measured accuracy, and how to test any tool
Before deployment the county set an accuracy target of 98%, and measured against calls marked by hand, the pipeline came in close to 99.5%. We set up that measurement with the county instead of offering a vendor estimate. The reports we turned in on the test sample listed every word spoken, when it was spoken, its PII class, whether detection caught it and whether it was redacted. The county verified half of those reports itself, and every one it checked was accurate.
A figure like 99.5% belongs to the calls it was measured on, because line quality, accents, local names and call scripts differ from one contact center to the next. What carries over is the method, and it's the one I'd use to test any call redaction tool, ours included. Work through four questions in order, because a failure at an early step can't be rescued by a later one.
- Every word has to be detected correctly rather than lost in background noise or a weak signal, since a word that never reaches the transcript can't be found by anything downstream.
- Each one then has to be recognized as the right language and after that as the right word, because a surname heard as a common noun gives detection nothing to catch.
- Whatever is flagged has to be real PII in the context of the whole conversation, so a date counts as a date of birth only when the caller is giving one.
- The redaction has to be complete, leaving nothing a human ear could pick up.
The last question is the easiest to overlook, because detection can be entirely right and the redaction can still give a word away. Take a caller who gives his name as Steve. If only the middle of the name is muted, the S and the V can survive at its edges, and a listener who knows him may still hear who it was. Redaction has to remove the audible sound beyond the points where the engine marked the word's start and end, which takes options in the redaction itself rather than one fixed cut. We built those options into Redactor, so each redacted segment can be padded at its start and end by an amount set to suit the recordings.
Word timings, confidence thresholds and the digits a transcript can drop before detection sees them are all laid out in finding spoken card numbers and names in a call.
Why we masked both channels, and the captions too
Every call at the county was recorded in stereo, with the agent on one channel and the caller on the other. It would have been easy to assume the caller's details stayed on the caller's side, but these conversations don't work that way. The caller gives a date of birth or an address, and the agent reads it back to confirm it. Often that's because the information is legally required to be confirmed, as most of it was for the county. Mask only the caller's channel and the same details stay audible in the agent's voice a few seconds later.
A flagged interval is therefore masked on both channels at once, which is how we built Redactor to treat stereo calls, even though the agent's words in that moment are lost too. Keeping the agent's side intact does not guarantee the PII is gone, and in a program whose redacted calls were leaving the county's hands, one read-back would have been a disclosure. In any pilot I'd include calls where the agent repeats a number, and play each channel alone at every redaction. Redacting stereo call recordings goes into how platforms lay out their channels, where per-channel transcription breaks down on real calls, and how to choose between a tone and silence.
The transcript and captions are the other place a masked detail survives, and they're easier to miss because nobody hears them. If the caption track sits outside the redaction job, the audio falls silent where the card number was while the text beside it still spells the number out. Redactor masks the captions in the same job and from the same detections as the audio. Redacting the text along with the audio follows every other copy a call can leave behind, down to the summaries AI tools write about it.
Two million files first, then every working day
The backlog came before the daily flow, so about two million stored recordings, as many as the county takes in over a year, went through the pipeline before the current calls did. The plan gave us 60 days from deployment to clear it and then go live on current calls. That meant scaling out at the start, with enough AI servers working the queue to get through the archive in time. Once it was done we scaled back in to the size of the daily flow. Because everything ran in the county's own AWS account, the county paid for every server that kept running, so shrinking afterwards was part of the plan from the beginning.
The daily flow is a different problem, because each day's calls have to clear within the day and small failures repeat thousands of times. Whichever stage runs slowest, whether transcription, detection, redaction or export, decides how many calls the pipeline can clear, and an average minutes-per-call figure hides which one that is. A failed job that never gives back its place in the queue can stop a day's calls instead of one file. A long call carrying more redactions than a pipeline step accepts can do the same, and so can a notification endpoint that goes offline. Sizing each stage and piloting on one real day of calls are the subject of planning bulk call recording redaction.
Reporting was one of the county's requirements, because at that volume a failure rate that looks negligible as a percentage is still a steady stream of calls. The county wanted to see how many calls were processed, how many failed and at which stage. It also wanted to be able to pull any recording, listen to what was redacted and correct it by hand. A call that never reached redaction is still sitting somewhere unredacted, so every failure has to be retried or dealt with, never quietly skipped. Until its redacted copy exists, access to a new recording should be limited to the people and systems that process it.
What to license, and where it runs
For everything in this guide you license Redactor, and the VIDIZMO platform it runs on comes with it, supplying the call library, users and permissions, storage connectors, export and notifications. It can be hosted by VIDIZMO, run in your own cloud account as it was for the county, installed on premises, or run air-gapped with detection on the platform's own AI. Redactor supports customers' PCI DSS obligations by detecting and redacting cardholder data from uploaded content.
If you redact calls on behalf of several client organizations, one deployment can serve them all with each client's content and users kept separate, or a client can have a deployment of its own. The platform supports white-labeling, and its REST API lets your own systems submit calls and collect the redacted results.
To see the whole flow on recorded calls, start with call recording redaction in Redactor.
TopicsRedactionAudio RedactionRedactor
About the author
Akhlaq Khan is Co-Founder and VP Delivery at VIDIZMO, where he oversees the full product portfolio including Redactor, AI Live Insight, and the company's AI processing pipeline. With over 20 years in software development and product management, Akhlaq leads the teams building VIDIZMO's AI-powered redaction engine, which automates PII, PHI, and PCI protection across video, audio, documents, and images for law enforcement, legal, and enterprise organizations. An AWS certified professional, he brings deep technical expertise in AI/ML workflows, compliance automation, and scalable SaaS architecture.
You may also like
Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe
I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...
Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera
A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...
Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video
A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...
See it on your own content
Tell us what you are trying to solve and we will show you how it works on your infrastructure.