Home / S3 to S3 pipeline
S3 to S3 Redaction Pipeline
No operator opens a file. No second copy of the archive. Redaction runs as a stage in a data pipeline rather than as a task someone performs.
When This Is the Right Shape
Three situations, all of them high volume and low judgement:
- Call-recording archives where PCI or PII has to come out before the recordings reach analytics, QA or a data warehouse
- Training data preparation, where PII must be removed before content is used to train or fine-tune a model
- Platform and BPO operators redacting on behalf of their own customers, where redaction is a processing step inside someone else's product
What these share is that the volume is too large for review and the decision is too routine to need it.
How It Runs
Watch or trigger
A new object lands in the source bucket, or your orchestration calls the API
Ingest
The file is pulled, not uploaded by hand
Detect and redact
Classes, masking style, confidence threshold and output policy are all set per job
Return
The redacted object is written to the destination bucket
Notify
A webhook reports completion against a documented event catalog
Because the output policy is configurable, the pipeline decides what happens to the source: keep it, recycle it, or redact in place.
No Duplicate Storage
The pattern that matters for cost at archive scale: content does not need to live permanently in a second platform. It is pulled, processed and written back. For organizations with petabyte archives and a retention obligation, that is the difference between a redaction step and a storage migration.
Throughput
Queue-based processing runs unattended, so a batch submitted at the end of a day is worked overnight. Volume has been exercised at roughly two million recordings a year.
Bulk redaction is permissioned per format, and API credentials carry the same per-format rights a user would.
FAQ
S3 to S3 Redaction Pipeline questions, answered
What does a bucket-to-bucket pipeline do?
Content is read from a source bucket, redacted, and written to a destination bucket, with no operator opening a file and no duplicate storage.
Can we scope which files are taken?
Yes. An extension allowlist decides which file types are imported, named folders are excluded, and the source hierarchy is preserved or flattened.
What happens to the source file?
It is left in place, moved, or deleted after ingestion, as configured. Because the output policy is configurable, the pipeline decides what happens to the original rather than assuming.
How do we know a job finished?
Webhooks fire against a documented event catalogue with delivery logs, so a downstream system reacts to events rather than polling for changes.
Does the output keep the original structure?
It can. Output is written back in the same structure and format, with accompanying files intact, which is what lets a downstream analytics or records system read it without changes.
Bring Us a Bucket
Point us at a sample set and we will run the pipeline end to end, including the webhook.