Redaction, Video Redaction, Redactor
Tracking Through Occlusion, Crowds and Scene Cuts in Video Redaction
Every tracked mask in a redacted video rests on a promise that the person under it on the ten-thousandth frame is the same person who was under it on the first. No frame in between should have gone uncovered either, and a reviewer checking tracks before a release is testing that promise over and over, across every person in the video. Tracking keeps it with little effort while one person walks across an empty room, and it comes under strain in the places our guide to video redaction best practices marks for the closest review.
Object tracking through occlusion fails in a different way from tracking across a cut or through a crowd, and each break has a different consequence for a redacted release. Knowing where tracking breaks tells a reviewer where to spend limited time, and it tells an evaluator which clips to test a tool on.
What a track promises, and three ways it breaks
A track links detections of one person across frames into a single identity, so that a decision made once, such as masking a bystander or leaving a suspect visible, holds for the whole recording. Where those detections come from is the subject of our article on object detection in video. The MOT16 multi-object tracking benchmark scores trackers on distinct errors "such as missed targets, ghost trajectories, or identity switches," and each of its error types maps onto a redaction risk.
An identity switch, in the benchmark's definition, is counted when a target is matched to a different track from the one it was last assigned to. In practice the track that was following one person has moved onto another. A fragmentation is counted "each time a trajectory changes its status from tracked to untracked and tracking of that same trajectory is resumed at a later point."
| Tracking failure | What happens on screen | What it means for a redacted release |
|---|---|---|
| Miss | The person is present and no box is | Exposed frames that anyone can pause on |
| Identity switch | The track jumps from one person to another | The decision to mask or show one person carries onto someone else |
| Fragmentation | One person's track breaks into several short tracks | More review work, and more places for an unmasked frame to hide |
Identity switches do the most damage when some people must stay visible, the case our article on selective video redaction is built around. There a switch can leave a bystander unmasked or hide the very person the release is about. On a body camera traffic stop, a switch as the officer's partner crosses in front of the driver can hand the driver's mask to the partner and leave the driver in plain view.
Ghost trajectories, the benchmark's term for tracks that follow nothing real, matter less in redaction because a mask over empty space exposes no one. They still cost review time, and a stray mask can hide something the requester is entitled to see.
Occlusion: the frame before and the frame after
A person who walks behind a parked van, a car door or another person should come out the other side with the same identity, so that every decision already made about them still applies. MOT16's annotators follow the same idea, marking each target "through occlusions as long as its extent and location can be determined accurately enough." A target that reappears only after a long absence, when its location during the occlusion is ambiguous, is given a new identity.
The edges of a gap need as much care as the gap itself, and a tracker that bridges the gap well can still get its edges wrong. One failure seen in the field treats the last detection before a gap as the end of the track and draws no mask on it. The frame a person was last seen on then goes out bare, a brief flash of the face at every occlusion. A related one misses the person reacquired near the end of a clip and seen only once more, leaving a single detection that nothing covers at all.
VIDIZMO Redactor treats a detection just before a gap as a real detection that needs its mask. Both edges of every gap therefore carry a mask, and so does a lone detection after the last gap. A person hidden for a long time can still come back as a new track in any tracker, and joining the two tracks is then a reviewer's job. How long a gap a tracker can bridge varies from tool to tool, so an evaluation clip should include an occlusion lasting several seconds as well as the brief ones.
Padding, which holds the last known box briefly past each end of a track, helps at exactly these boundaries. Our article on frame-by-frame redaction explains its trade-off with fast subjects, and what goes wrong when padding reaches every frame of a track instead of its two ends.
Scene cuts, flashes and edited interviews
A hard cut to a different shot should reset tracking, so that an identity never carries from one scene into another where a different person happens to stand in the same place. Raw body camera and closed-circuit television (CCTV) recordings seldom contain hard cuts, so most of the cuts a reviewer meets are in edited material such as interviews, compilations and exports that switch between cameras. The cost of resetting is that the same person on both sides of a cut becomes two tracks, which a reviewer merges when the release needs one decision for both.
Deciding where a cut falls takes more care than it seems to, and two failures seen in the field show why. A measurement that counts small brightness changes as change reads every frame as a hard cut, so tracking resets constantly and every track shatters into fragments, multiplying review work without any single frame looking wrong. The black bars of letterboxed video cause the opposite trouble, since they never change and so dilute any measurement that includes them, which makes cut detection depend on the shape of the footage. Redactor judges a cut on the picture area alone.
A camera flash, a strobe or a sudden exposure jump changes the whole picture without changing the scene, and a tracker may read the change as a cut. Footage full of them needs a closer look at the frames just after each one. Edited interviews need the same care, because three people filmed one after another in the same chair and the same framing can end up joined into one face track. That matters most when only one of them may be shown.
Crowds and dense traffic
Faces in a crowd are small, turned away and partly hidden by other faces, and neighbors move together, so every difficulty above arrives at once. The authors of the WIDER FACE benchmark describe the faces in their dataset as "extremely challenging due to large variations in scale, pose and occlusion." They also report "a gap between current face detection performance and the real world requirements."
In a crowd a tracker can fuse two neighbors into one track, or drop a face that stays hidden for most of a segment, and neither failure is obvious at normal playback speed. A reviewer who has to ration time should spend it on those segments first, even when the rest of a recording can be checked at normal speed. A fused track often gives itself away when the two people separate, because the box jumps from one to the other or stretches to cover both. The frames where a group breaks up are therefore the ones to step through.
Busy traffic produces a related problem with license plates, which are small, fast and often readable for only a few frames. Even when every plate is masked, each one can break into many short tracks, and every fragment is one more item for a reviewer to confirm. When the camera itself is moving, tracks break more often and for more reasons, which our article on redacting footage from a moving camera covers device by device.
Where the reviewer takes over, and what Redactor gives them
Every break described above ends with a reviewer, so the tools for correcting tracks decide how long a release takes. Easy footage tells an evaluator almost nothing about where a tool will fail, which is why a useful test set is small and deliberately hard. Pick a clip where people pass behind vehicles and each other, an edited sequence with several cuts, a crowded scene and a stretch of busy traffic. Run the tool, step through the moments listed in the checks below, and count what a reviewer had to fix in each clip. Time the corrections as well, since minutes of review per minute of footage is the figure that decides how long a unit's releases will take.
A reviewer working through those failures needs operations that act on whole tracks rather than single frames. Redactor's review studio shows each detection as a track across the timeline, so every correction applies to the person's whole appearance:
- Merge two tracks that belong to one person, such as the two sides of a long occlusion or of a cut.
- Split a track at the frame where it jumped to someone else.
- Adjust a box that has drifted, rename a track, or delete a false detection.
- Draw a missed person once and let tracking carry the box forward.
Tracking a hand-drawn box usually runs in the reviewer's browser, so the box follows the person as soon as it is drawn. When the browser cannot do it, the same request runs on the server. The detections that drive the masking are also indexed, so an investigator can search a library for people, vehicles or license plates and land on the moment each one appears. Release is operator-gated, with the reviewer correcting and adding to the tracks before any output is produced.
Checks before a tracked video goes out
These checks take a reviewer straight to the places where tracks break, and they work the same way whichever software produced the tracks.
| Where to look | What to confirm |
|---|---|
| Every moment someone passes behind something | The frame before and the frame after the gap are both masked |
| The opening frames after each cut or flash | Masks restart on the right people |
| The unmasked person in selective redaction | It is the same person from the first frame to the last |
| Crowded segments and dense traffic | No face was dropped and no two people share a track |
For the same rules applied to faces from detection through review to release, see how Redactor redacts faces in video.
TopicsRedactionVideo RedactionRedactor
About the author
Nabeel Ali is a Senior Product Analyst at VIDIZMO, with five years at the company. An engineer by profession, Nabeel works on audio, video and DICOM redaction in Redactor, and has also worked on AI Intelligence Hub and AI Live Insight. The role sits where customers and engineering meet: hearing from customers, customer success, and the sales and marketing teams what goes wrong in real deployments, then working with the engineers until the product answers it.
You may also like
Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe
I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...
Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera
A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...
Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video
A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...
See it on your own content
Tell us what you are trying to solve and we will show you how it works on your infrastructure.