Redaction, Audio Redaction, Redactor

Redacting Stereo Call Recordings: Two Channels, One Clean Result

A QA analyst who puts on headphones to check a redacted call can hear the customer's card number silenced on one side and the agent reading it back, clearly, on the other. That is the characteristic failure of stereo call recording redaction, and it comes from treating a two-channel file as if each channel held only one person's speech. Stereo recordings are common in contact centers, and they carry PII on both channels for reasons that have little to do with who was speaking.

Stereo handling is also one of the easier failures to test for, and our guide to call recording redaction at scale puts it among the checks every pilot should run. The test only tells you something, though, if whoever runs it knows how the channels are laid out on their own platform.

How platforms lay out a stereo call

No single convention decides which side of a conversation lands on which channel, even among the largest platforms. Amazon Connect stores agents and contacts on separate stereo channels, according to its recording documentation. For agent interactions the agent is on the right, and the customer shares the left channel with any conferenced third party. For automated IVR segments the layout changes, with the customer on the right and the system prompts on the left.

Telephony platforms commonly offer both mono and dual-channel recording, and the difference matters to anyone redacting the result. In a mono recording both legs of the call are mixed into a single channel, while a dual-channel recording gives each leg its own channel within one file. Which channel holds the agent can depend on how the call flow was built, so it has to be confirmed for each platform and each flow rather than assumed. A platform that mixes both legs down before storage removes the channel question altogether and replaces it with overlapping speech in a single track, which is harder to attribute and no easier to redact.

Recordings from older phone systems are often single-channel G.711, the telephony codec that the ITU-T recommendation defines at a nominal 8,000 samples per second. Each sample uses eight binary digits under one of two encoding laws, A-law or µ-law. Audio of that kind, or GSM 6.10, often sits inside a file with a .wav extension, and some audio decoders cannot decode such files reliably or seek within them. A tool that assumes a .wav extension means plain PCM will stumble on exactly the oldest part of an archive. A stereo test set should therefore include a handful of those recordings as well, so that the codec question is answered before production rather than during it.

Where PII sits in a two-channel call

Sensitive data rarely stays on the channel of the person it belongs to, and the reasons are entirely ordinary. The customer says the card number, and the agent confirms it a few seconds later by reading it back, which puts the same digits on the other channel. An IVR prompt may ask for the number before an agent joins, and on some platforms the prompt and the caller sit on different channels of that segment. Headset bleed and cross-talk carry a quieter copy of each voice onto the opposite channel, and on Connect a conferenced third party shares the customer's side.

Each of those paths defeats a redaction that masks only the channel where it believes the speaker was. The read-back leaves the number audible in the agent's voice, bleed leaves a faint copy on the other side, and a conference call puts a second voice on a channel where the tool expected one. A faint copy is still a disclosure, since turning up the volume on one channel is the first thing a curious listener does. How the number is found in the first place, including why digits said in groups slip past detection, is the subject of our article on finding spoken PII in a call.

Masking the flagged interval on every channel

The dependable way to close all of those paths is to treat an interval flagged anywhere in the call as flagged on every channel, and to silence or beep it on all of them at once. Attributing each detection to a single channel would reproduce the read-back, bleed and conference leaks on every call that contains them. At contact center volume that means a steady trickle of disclosures that nobody hears until a complaint arrives.

Masking every channel means the other party's words in the same interval are lost too, which errs toward privacy. The loss is usually small, because flagged intervals are short and a listener rarely needs the half-second of agent speech that overlapped a card number. A QA program that scores agent behavior should still know about the trade-off before it starts marking agents down for words it can no longer hear.

The redacted file should keep the original's channel count, sample rate and duration, so that QA tools and analytics platforms accept it as they accepted the source. For uncompressed PCM WAV, a redaction can go further and leave every byte outside the masked intervals untouched, as our article on byte-level redaction and chain of custody explains.

Per-channel transcription: what it helps and where it fails

Transcribing each channel separately is common across the speech industry, and it makes the question of who said what much easier to answer. Amazon Transcribe's channel identification transcribes the speech on each channel separately and does not support audio with more than two channels. Its Call Analytics mode only supports two-channel audio, with an agent on one channel and a customer on the other.

Per-channel processing also brings failure modes, and they show up in exactly the recordings a contact center produces. When both parties talk at once, AWS notes that the timestamps on the two channels overlap, and bleed can put fragments of one speaker into the other channel's transcript. Hold music or recorded prompts on one channel produce text that nobody said to the customer. Very different levels between channels can push the quieter side below what a recognizer handles well. Channels encoded separately can also drift apart in time, so an interval found on one no longer lines up with the same moment on the other.

Channel-based attribution also differs from full speaker separation, because a channel can hold more than one voice, as any conference call shows. Redacting only one party's speech, such as only the customer's, depends on every one of these conditions holding at once. A vendor offering per-channel or per-speaker redaction should be able to answer these questions on your own recordings:

  • What happens to a read-back, where the customer's number is repeated in the agent's voice on the other channel?
  • Can a faint copy of a number survive through bleed on a channel that was left unmasked?
  • How are hold music, IVR prompts and very different channel levels treated?
  • What happens on a conference call, where a third voice shares a channel?
  • Does a redacted interval stay aligned across channels that drift apart?

Tone or silence on a stereo call

The choice between a tone and silence looks cosmetic until a QA listener or a court has to interpret the result. SWGDE's video and audio redaction guidelines, version 2.3 from December 2025, recommend a single tone or set of tones as the redaction signal. They treat silence as generally unsuitable, because a redacted segment can be confused with original content that contains silence. Silence is reserved for cases where several simultaneous channels need independent redaction and an inserted tone risks masking content on another channel.

Amazon Connect's own analytics takes the other path and redacts sensitive data in its audio files as silence. Its redaction documentation notes that those silent stretches are not flagged anywhere as non-talk time, which is the confusion SWGDE warns about.

A tone tells a QA listener exactly where something was removed, while silence avoids a sound that could be mistaken for part of the call. On a stereo call masked on every channel for the same interval, SWGDE's concern about a tone covering content on another channel carries less weight, because that content is being masked too. What matters most is choosing per use and recording the choice, so that whoever hears the file later knows what a silence or a tone means.

Checks to run on stereo output

A short test set of your own calls will show whether a tool handles stereo correctly, and none of these checks needs special equipment. Choose calls with read-backs, transfers and at least one conference, because those are the recordings where channel handling goes wrong.

Test Method Pass
Each channel alone Play the left and right channels separately at every redaction Silence or tone on both channels for the same interval
A read-back call Redact a call in which the agent repeats the customer's card number No digit is audible in either voice
A three-way conference Redact a call with a third party on the customer's channel The third party's PII is masked like the customer's
Bleed at volume Raise the level of each channel at every redaction Nothing intelligible survives on either side

What Redactor does with two channels

VIDIZMO Redactor masks a flagged interval on every channel of a stereo or multi-channel recording, with silence or a beep chosen per job, so a read-back or a bleed-through copy is covered wherever it lands. Detection runs across the whole recording rather than channel by channel, which puts both voices in the transcript it reads. A PCM WAV call redacted for detected spoken PII keeps its sample rate, bit depth, channel layout and duration, so a QA tool receives the stereo file in exactly the format it was recorded in.

For the wider audio workflow, from transcript search to bulk masking, see Redactor's audio redaction software.

TopicsRedactionAudio RedactionRedactor

You may also like

Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe

I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...

Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera

A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...

Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video

A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...

See all blogs

See it on your own content

Tell us what you are trying to solve and we will show you how it works on your infrastructure.