Redaction, Video Redaction, Redactor
Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video
A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, stored apart from the frames and sometimes more than once. When the summary and the frames disagree, a variable frame rate video can play perfectly while redaction software that trusts the summary drifts. That drift is one of the failures our guide to video redaction best practices traces across the length of a recording.
A drifting mask on a long recording is easy to mistake for a tracking failure, and the two need different fixes. A blur that sits squarely on a witness's face for the first ten minutes of an interview, then covers the wall beside her an hour later, may never have lost track of her. The software may instead have placed each mask by counting frames at the declared rate, in a file whose frames did not run at that rate. The same mistake explains a redacted file that ends while its audio keeps playing.
What variable frame rate means inside a video file
Apple's QuickTime File Format specification states the model directly: "Samples are stored in the media, and they may have varying durations." For video a sample is a frame, and a table called the time-to-sample atom does the bookkeeping, "providing a mapping from a time in a media to the corresponding data sample." The frame rate a player shows is worked out from that timing rather than the other way round.
The MP4 files that records units handle every day inherit that model of per-frame timing. The Library of Congress lists MP4 as a subtype of the ISO Base Media File Format, and it notes that MP4's versions "owe a debt to the QuickTime file format that preceded them."
In a constant frame rate file every frame lasts the same length of time, so the summary and the frames agree and simple arithmetic works. In a variable frame rate file the durations differ from frame to frame, so the single frames-per-second figure is at best an average. A file can also store more than one version of that figure, typically a base rate and an average rate, and nothing forces the versions to agree with each other or with the frames.
A mismatch between header and frames is uncommon, but it turns up in ordinary evidence files and not only in damaged ones. A sweep of 597 real recordings from our test corpus found six such files, all from body cameras, dash cams or closed-circuit television (CCTV) systems. Their frame counts differed from what their headers implied by a factor of 1.8 to 3.3. Six in 597 is about one file in a hundred, a small minority. An agency that releases hundreds of recordings a year will still meet some of them, and nothing in a file's stated frame rate marks which ones they are. The warning sign is a mismatch between numbers: when a media inspector's count of frames sits far from the duration multiplied by the stated rate, something in the file is not describing the frames.
Two frame-rate fields, two different answers
The clearest examples in our corpus were two recordings in which a different field was wrong in each. A 13.5-hour CCTV recording declared an average rate of 60 frames per second while running at about 29.97, and its base-rate field was right. A 41-minute body camera clip declared a base rate of 60 while running at about 30, and its average-rate field was right. We established the true rates from the frames themselves, since no field in either file could be taken on trust.
Software that trusted the average rate would have been wrong about the first file, and software that trusted the base rate would have been wrong about the second. Masks placed by counting frames at a doubled rate drift further from the face with every minute of playback. A file encoded at the doubled rate plays in half the time of its audio, which on the CCTV recording would put picture and sound hours apart by the end.
The only ground truth in either file was the timing of each frame, which is why every mask in VIDIZMO Redactor follows its own frame's presentation timestamp. A header can then be wrong in either field without moving a single mask, because no mask position is ever computed from it.
Placement is only half of the problem, since the redacted picture still has to be written out at some rate, and a wrong header can mislead that choice as well. Whichever tool produced the file, the sound checks at the end of this article are the way to catch it.
Durations, rounding and frame counts
A container's overall duration generally reflects its longest stream, so when the audio outlasts the video, a frame count computed as duration times rate includes frames that do not exist. We have seen a 74-second video carrying a 135-second audio track, and a far larger case in a body camera file that our article on redacting moving-camera footage describes. The rule is to count the frames the file really contains and to confirm the count finished, because a list of frames cut short partway would otherwise pass for the whole video.
Rounding creates a quieter version of the same problem when a frame rate passes from one system to another. A source running at 1000/33 frames per second, which is 30.303, becomes 30.3 wherever something along the way keeps only one decimal place. That rounding loses one frame in every ten thousand, which over a three-hour interview puts the picture about a second out of step with the sound. Redactor therefore reads the rate from the file itself instead of accepting a rounded value from elsewhere.
Start times raise a related issue in variable frame rate media, where rounding can give two consecutive frames the same start time. Software that identifies frames by start time alone lets those two frames share one answer, so one of them carries the other's masks or none at all. Each frame has to be told apart by more than its rounded start, and Redactor keeps the masks of two such frames separate. Timing matters again for the frames between analyzed ones, which our article on frame-by-frame redaction covers.
Keeping sound and picture together
A mask can sit on exactly the right frame of a file whose picture and sound disagree, so timing has an audio side as well as a visual one. The Scientific Working Group on Digital Evidence (SWGDE) sets out the audio side in its Video and Audio Redaction Guidelines. It asks for the redacted recording to be exported "with the original source properties," so that display resolution and frame rate match the source. The guidelines also warn that when audio is redacted, "care must be taken that the audio and video streams remain synchronous."
A tone that bleeps a spoken name has to land on the same moment as the speaker's lips, or the redaction either misses the word or cuts into the words around it. Our article on redacting audio recordings covers the audio half of that job, from finding the spoken details to replacing them.
SWGDE's workflow also asks the practitioner to interrogate the recording for its "display resolution, pixel aspect ratio, frame rate and codec and audio sampling rate" before redacting. That is the right instinct, provided the frame rate found there is treated as a claim to verify. Proprietary recorder formats add a step before any of this, since the same guidelines note that transcoding "may be necessary to allow a recording to be redacted." Every conversion is one more chance for a file's timing to change.
What Redactor does with a variable frame rate source
Alongside the timestamp rule, Redactor converts a variable frame rate source to a constant frame rate before redaction, so masking works on frames of equal length. The redacted copy is therefore constant frame rate and does not reproduce the source's variable timing, and the conversion is an extra encode that adds processing time on long recordings. Whatever the source container, the redacted copy is also a new H.264 MP4 file. Both the constant frame rate and the new format therefore belong in the redaction record, a point our guide makes for copies that may be played in court.
Checks you can run on any redacted file
None of these checks needs access to the redaction software, only the original file, the redacted file and a media player that can step one frame at a time.
| Check | How to run it | What a failure looks like |
|---|---|---|
| Duration and frame count | Compare the redacted file's duration and frame count with the original's in a media inspector | A redacted file noticeably shorter than the original, or missing a large share of its frames |
| Mask position late in the file | Scrub to the last part of a long recording and step through a few seconds where a masked face is visible | The mask sits beside the face instead of on it |
| Lip sync late in the file | Watch a speaker near the end of the recording, and listen to any redaction tones | Lips and sound out of step, or tones landing early or late |
| Timecodes in the redaction log | Record positions by the file's runtime, and note any offset between the burned-in clock and real time | Log entries that point at the wrong moment |
A variable frame rate source converted to a constant rate can legitimately end up with a slightly different frame count, so treat the duration as the firmer of the two numbers for those files. During an evaluation, run the checks on files your agency actually receives, since clean sample clips will not show how a tool copes with a header that is wrong.
Burned-in clocks need separate handling when you write timecodes into a redaction log, because the clock on a recorder can be wrong. SWGDE's best practices for acquiring video from digital video recorders (DVRs) tell examiners to "calculate if there is a time offset between real time and the DVR system clock." They also tell examiners not to change the time and date on the recorder itself. The redaction guidelines add that embedded timestamps "may not be as useful a reference as file runtime," so log positions by runtime and note the offset separately.
Clipping a segment for release raises the same timing questions at the start and end of the clip. Our article on clipping CCTV footage for a subject access request covers choosing a clip window and confirming that timestamps still line up afterward.
Long surveillance recordings are where a wrong header does the most damage, and CCTV redaction in Redactor follows the same timestamp rule.
TopicsRedactionVideo RedactionRedactor
About the author
Nabeel Ali is a Senior Product Analyst at VIDIZMO, with five years at the company. An engineer by profession, Nabeel works on audio, video and DICOM redaction in Redactor, and has also worked on AI Intelligence Hub and AI Live Insight. The role sits where customers and engineering meet: hearing from customers, customer success, and the sales and marketing teams what goes wrong in real deployments, then working with the engineers until the product answers it.
You may also like
Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe
I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...
Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera
A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...
Redacting 360-Degree Video and Images: Distortion, Seams and Playback
Open a 360-degree photo in an ordinary image viewer and it looks like a warped panorama, with the ceiling smeared along ...
See it on your own content
Tell us what you are trying to solve and we will show you how it works on your infrastructure.