Home / S3 to S3 pipeline

S3 to S3 Redaction Pipeline

No operator opens a file. No second copy of the archive. Redaction runs as a stage in a data pipeline rather than as a task someone performs.

When This Is the Right Shape

Three situations, all of them high volume and low judgement:

  • Call-recording archives where PCI or PII has to come out before the recordings reach analytics, QA or a data warehouse
  • Training data preparation, where PII must be removed before content is used to train or fine-tune a model
  • Platform and BPO operators redacting on behalf of their own customers, where redaction is a processing step inside someone else's product

What these share is that the volume is too large for review and the decision is too routine to need it.

How It Runs

01

Watch or trigger

A new object lands in the source bucket, or your orchestration calls the API

02

Ingest

The file is pulled, not uploaded by hand

03

Detect and redact

Classes, masking style, confidence threshold and output policy are all set per job

04

Return

The redacted object is written to the destination bucket

05

Notify

A webhook reports completion against a documented event catalog

Because the output policy is configurable, the pipeline decides what happens to the source: keep it, recycle it, or redact in place.

No Duplicate Storage

The pattern that matters for cost at archive scale: content does not need to live permanently in a second platform. It is pulled, processed and written back. For organizations with petabyte archives and a retention obligation, that is the difference between a redaction step and a storage migration.

Throughput

Queue-based processing runs unattended, so a batch submitted at the end of a day is worked overnight. Volume has been exercised at roughly two million recordings a year.

Bulk redaction is permissioned per format, and API credentials carry the same per-format rights a user would.

FAQ

S3 to S3 Redaction Pipeline questions, answered

What does a bucket-to-bucket pipeline do?

Content is read from a source bucket, redacted, and written to a destination bucket, with no operator opening a file and no duplicate storage.

Can we scope which files are taken?

Yes. An extension allowlist decides which file types are imported, named folders are excluded, and the source hierarchy is preserved or flattened.

What happens to the source file?

It is left in place, moved, or deleted after ingestion, as configured. Because the output policy is configurable, the pipeline decides what happens to the original rather than assuming.

How do we know a job finished?

Webhooks fire against a documented event catalogue with delivery logs, so a downstream system reacts to events rather than polling for changes.

Does the output keep the original structure?

It can. Output is written back in the same structure and format, with accompanying files intact, which is what lets a downstream analytics or records system read it without changes.

Bring Us a Bucket

Point us at a sample set and we will run the pipeline end to end, including the webhook.