Redaction, Audio Redaction, Redactor
Bulk Call Recording Redaction: Planning for Thousands of Calls a Day
Four questions decide whether a bulk call recording redaction plan survives its first busy week. They ask how many hours of audio the pipeline must clear each day and which stage sets the pace. They also ask what happens when a job or a downstream system fails, and what becomes of the original recordings once their redacted copies exist.
A public-sector contact center once put the first of them to us directly. It asked how many servers it would need if it had to redact tens of thousands of short calls a day, each within a day of recording. That question turns bulk redaction from a feature comparison into capacity planning, and it has no answer until each stage of the pipeline is measured on its own. Volume is where several failures in our guide to call recording redaction stop being theoretical, and planning for it starts with arithmetic.
How many hours a day, and which stage sets the pace
Daily hours of audio come first, calculated as the number of calls multiplied by their average length. As a worked example with invented round numbers, 12,000 calls a day at an average of five minutes each comes to 1,000 hours of audio. Every one of those hours passes through several stages, and the stages have very different costs.
The time budget per call comes next, and the obligation the redaction serves sets it. Payment calls may need to move faster than the rest. The PCI SSC expects a card verification code that reaches a recording to be removed as soon as the transaction is authorized, as our guide explains.
Each stage then needs its own throughput figure, measured on your own recordings and hardware rather than taken from a brochure. Transcription, detection, redaction and export move at different speeds, and a day's audio has to clear the slowest of them inside the time budget, whatever the others can do. Where transcription runs in a cloud account, the provider's cap on concurrent jobs belongs in the plan beside the server count, along with the time it takes to have that cap raised.
We answered the contact center's question by timing transcription, detection and redaction separately, documenting the cloud transcription service's parallel capacity and how to raise it, and sizing from the slowest stage. An average minutes-per-call figure would have hidden the stage that actually constrained the night.
Why short calls cost more than their length
Every file carries fixed costs that have nothing to do with how long it is, and a contact center's calls are mostly short. Transferring the file, setting up each job and encoding a playback copy take much the same time for a two-minute call as for a twenty-minute one. On short calls those fixed costs can outweigh the audio work itself.
The practical response is to skip every step that the workflow does not actually need for a given kind of recording. Audio that is only redacted and exported onward, and never played in a browser, may not need a playback encoding at all. Choosing processing per upload turns that into a configuration decision. The trade-off is that content left untranscoded stays in its original format, which a browser may not play, so the choice belongs to whoever owns the downstream use.
Queued, unattended and fed from storage
A day's calls can only be cleared overnight if the work runs as a queue that nobody watches, so that the morning starts from results rather than from a backlog. Throughput should grow by adding AI servers to the queue without reconfiguring the application, which turns capacity into a hardware decision that can be tested during the pilot.
A queue is only as steady as its intake, and at this volume a call that is ingested twice, or never, becomes a daily reconciliation job for someone. Getting each call into the queue exactly once, with its metadata, and back out again in the layout the next system expects is covered in our article on Amazon Connect recordings in S3.
When jobs and endpoints fail
Most failures that matter at volume are small ones that behave differently when they repeat thousands of times, so the questions to ask a vendor are about how a job ends rather than how it runs. A job that fails early, because its content was deleted after it was queued or a validation check failed, should reach a final failed state with its original error kept. If it stays marked as running instead, it goes on counting against whatever limit applies to that kind of work. Enough such jobs stop new work from starting anywhere on the platform, and where one queue serves every format, a stall that starts with documents holds up call recordings too.
The longest calls are the ones most likely to hit a limit nobody planned for. A long recording dense with PII can produce more redacted spans than a step in some pipelines accepts, and a job that cannot run at all leaves exactly the calls with the most sensitive content unredacted. Ask whether the number of redactions in one recording has a ceiling, and put the longest, densest call you have into the pilot.
Downstream systems fail as well, and a notification endpoint that goes offline shows whether the queue depends on it. When delivery retries in a tight loop, the backlog of messages can hold up processing until someone restarts a service. Every failed delivery should be recorded, retries should back off to a limit, and delivery should never block the redaction work itself. Some faults are invisible at ten requests and serious at thousands, such as a small resource left behind by every request, which is why soak tests belong at production volume and for production durations.
Beneath those cases sits the fail-closed rule that our article on audio redaction and chain of custody describes for any tool. NIST SP 800-53 offers vocabulary for writing it into a requirement, even for systems its controls do not formally govern. Its SC-24 control, Fail in Known State, says failure in a known state prevents the loss of confidentiality, integrity, or availability of information when systems fail.
The original at volume
Keeping every original beside its redacted copy at least doubles the audio storage, which turns a custody decision into a capacity one once the calls number in the thousands each day. Recycling the original moves the decision to the recycle bin's retention period, and overwriting it removes the unredacted record for good. For payment calls the choice narrows further, since PCI DSS leaves little room for an unredacted original that still holds a security code. Our article on chain of custody for redacted calls takes up that point alongside the case for keeping originals elsewhere.
Retention schedules usually settle the choice for other calls, because they say how long the unredacted recording must or may be kept. An unredacted original kept only because nobody set a policy is what a written retention period prevents, and at contact center volume the backlog of such originals grows fast.
A pilot to run before committing
The pilot that answers the sizing question is one real day of calls run end to end, from ingest to export, on hardware and settings configured like production. Detection quality is a separate measurement that no amount of capacity fixes, and our article on audio PII redaction sets out how to test it.
| What to measure | How to measure it | What it tells you |
|---|---|---|
| Items at each stage, every hour | Count ingested, transcribed, redacted and exported items hourly | Where the backlog forms |
| Stuck items | List anything that has sat in one state longer than expected | Failures that are holding capacity |
| Stage timings | Time transcription, detection, redaction and export separately | Which stage sets the pace |
| Export status | Track each item's export state at the destination | Whether anything is lost or sent twice |
| Queue recovery | Inject a corrupt file and cut a downstream endpoint mid-run, then time the return to normal throughput | Whether one failure costs minutes or the night |
Scale the estimate from the slowest stage, then add headroom for the busiest day of the week and for the cloud quota you have actually been granted. An archive backlog is a different problem from the daily flow, with different economics, and it should be sized separately.
Running a day of calls through Redactor
VIDIZMO Redactor runs bulk redaction as queued, unattended work, permissioned separately for audio, video, documents and images, and it has been exercised at over 1.1 million recordings. AI processing scales out by adding servers to the queue with no application reconfiguration, and each file is redacted as its own job, so a failure affects one file rather than the batch. One recording has no ceiling on the number of redactions it can carry. Processing is chosen per upload, each redaction run has a configurable time limit, and a failed or timed-out job leaves the original intact and deletes its partial output. Webhook deliveries are logged with their outcome and retries, and the retry count and delay are configurable.
Service providers that redact recordings for several clients can run them from one deployment, with each client's content and users kept separate, or give a client a dedicated deployment where its contract requires one. The whole flow can also be driven through the REST API, so a provider's own systems submit calls and collect redacted results without anyone opening the interface. None of this needs more than Redactor and the platform it comes with, as our guide to call recording redaction sets out.
To run the same flow headless from one bucket to another, see bulk S3 to S3 redaction.
TopicsRedactionAudio RedactionRedactor
About the author
Nabeel Ali is a Senior Product Analyst at VIDIZMO, with five years at the company. An engineer by profession, Nabeel works on audio, video and DICOM redaction in Redactor, and has also worked on AI Intelligence Hub and AI Live Insight. The role sits where customers and engineering meet: hearing from customers, customer success, and the sales and marketing teams what goes wrong in real deployments, then working with the engineers until the product answers it.
You may also like
Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe
I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...
Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera
A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...
Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video
A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...
See it on your own content
Tell us what you are trying to solve and we will show you how it works on your infrastructure.