PHI Redaction, HIPAA Compliance, Redactor

De-identifying Multi-Frame DICOM Files, Series and Studies

A single magnetic resonance (MR) visit can reach a trial's central reader as one enhanced file holding hundreds of frames, or as hundreds of single-slice files spread across several series. Either way it is one study, and a de-identification process has to handle a multi-frame DICOM (Digital Imaging and Communications in Medicine) file as carefully as a folder of single-slice files.

Each shape carries its own kind of risk, and the checks that catch one can miss the other. Inside a file, the danger is missing frames or misreading how they are laid out. Across files, the danger is that de-identified pieces no longer fit together, or fit together so well that they lead back to the patient. Our guide to DICOM de-identification covers the header and pixel surfaces in general, while the detail here follows imaging from the single file up to the whole study.

What a multi-frame DICOM file is, and how series and studies differ

A DICOM instance is a single object, usually stored as one file, and a multi-frame instance is one whose pixel data holds many frames, with the count given by Number of Frames (0028,0008). Enhanced computed tomography (CT) and MR objects, ultrasound cine loops and multi-frame secondary captures all work this way.

A series groups the instances produced by one acquisition or reconstruction, and every instance in it carries the same Series Instance UID, a unique identifier (UID) for the series. A study groups the series from one imaging exam under a Study Instance UID, so a single visit can produce one study with several series in it. When studies are written to exchange media, a DICOMDIR file indexes them, and its patient records carry Patient's Name and Patient ID.

Level What it is What groups it What can go wrong
Frame One image inside an instance Its position in the pixel data A skipped frame keeps its burned-in text
Instance One DICOM object, single-frame or multi-frame Its own instance UID Header and pixels both need cleaning
Series Instances from one acquisition or reconstruction A shared Series Instance UID Files that stop sharing it stop grouping
Study Series from one imaging exam A shared Study Instance UID Inconsistent replacement splits or merges studies
DICOMDIR An index of the files on exported media Directory records Its patient records name the patient

Every frame, laid out the way the file says

Inside a multi-frame file, every frame has to be found and processed, and the layout of the pixel data has to come from what the file declares about itself. A block of pixel data with three planes shows why, because it can be three grayscale frames or one color image with red, green and blue samples. A tool that decides from the shape of the data can apply the mask along the wrong axis, and burned-in text then survives with no error raised. We read Number of Frames and Samples per Pixel before touching a pixel, since those two attributes settle which reading is right.

The masking style has to match on every frame, so that a cine loop does not flicker between a solid bar and a blur as it plays. The frame count and order must not change either, because a reader scrolling through expects the sequence the scanner produced. What those frames tend to carry, from ultrasound banners to captured screens, is the subject of our article on burned-in text in ultrasound and secondary captures.

A check we recommend when evaluating any tool uses a short CT series of 20 frames with different text on different frames. Put a name and a record number on frames 1 to 5, a name and a date of birth on frames 6 to 10, and no text at all on the remaining ten. After de-identification, all 20 frames should be present in the original order, the text should be masked on the first ten, and the last ten should be unaltered pixel for pixel.

When the regions to mask and the frames disagree

In any pipeline that finds text first and masks it afterwards, a region can end up addressed to a frame position the file does not contain. When that happens nothing is masked, and the file comes out exactly as it went in. A tool that counts the regions it was asked to mask, rather than the regions it actually masked, will then release that file labeled as redacted.

A file that nothing was removed from should never come back labeled redacted, and VIDIZMO Redactor treats that situation as a failed job. It counts the regions it actually masked, and a job that masked none of the regions it was given fails and releases nothing, so a silent failure becomes one an operator can see.

Header cleaning inside enhanced multi-frame files

Enhanced multi-frame objects move much of their header into sequences, which changes where a de-identification pass has to look. PS3.3 keeps attributes shared by all frames in the Shared Functional Groups Sequence (5200,9229) and those that vary in the Per-Frame Functional Groups Sequence (5200,9230). For the per-frame sequence it adds, "The first Item corresponds with the first Frame, and so on", so a file with hundreds of frames carries hundreds of items. A header script that only reads top-level attributes never opens those items, and anything identifying inside them leaves with the file.

Functional group items are nested sequences like any other, so they need the same depth of search as the rest of the header. Our article on private tags and nested sequences walks through how sequences nest and how to check them.

Keeping a de-identified study coherent across files

Across files, the de-identified pieces have to stay consistent with each other, and in many imaging cores that is the work of the honest broker's crosswalk and the header pipeline that applies it. For a trial, that consistency decides whether a central reader can use what arrives. Picture a study exported as a folder of several hundred single-slice files with an index file beside them. De-identified one by one with no shared crosswalk, the pieces can stop grouping into series, collide with other studies on import, and leave the patient's name in the index.

The same patient has to receive the same pseudonym in every file, or one person becomes several in the recipient's system. Replacement UIDs have to be consistent as well, because UIDs group instances into series and link derived images to their sources. The Medical Image De-Identification (MIDI) Task Group's report says that "within a defined scope of referential integrity, UIDs must be replaced consistently". That scope, it adds, "must at least include all objects within a single DICOM study, and preferably all objects for a single patient that may have multiple studies".

The report describes two techniques in general use, a persistent map from original to replacement UIDs and a one-way, deterministic hash of the original, and each has a cost. A map has to survive for as long as new objects may arrive, and every process working on the study has to share it. That need to share is the reason to keep the map in one pipeline under the organization's control, and a map kept for the long term must also be securely protected. A hash avoids the map, but the report notes the long-term risk that the cryptography behind it is eventually broken, while publicly shared collections may live forever.

A pseudonym that lets the source re-identify a patient later is a re-identification code under the Health Insurance Portability and Accountability Act (HIPAA). 45 CFR 164.514(c) allows one only if the code is "not derived from or related to information about the individual" and the mechanism for re-identification is not disclosed. A pseudonym built from the medical record number is derived from information about the individual, and a map shipped with the data discloses the mechanism, so both break the rule.

Dates raise the same consistency question whenever the intervals between a patient's scans matter to the analysis. PS3.15's Modified Dates option asks for dates to be changed in a way that "preserves the gross longitudinal temporal relationships between images obtained on different dates". The standard's note on shifting dates says the precise intervals survive only when the whole set is de-identified at the same time. The alternative is a mapping or database kept "to repeat this process on separate occasions", which raises the same custody question as a pseudonym map. Legal replacement values for dates, and for every other data type, are tabled with the DICOM tags that contain PHI.

Exported media add one more file to think about, because a DICOMDIR's patient records carry Patient's Name and Patient ID. PS3.15 is direct about it, stating that "Any existing non-de-identified DICOMDIR File shall be removed from the File-set."

Checks at the level of the study

Whoever owns the crosswalk, these checks show whether a de-identified study still works as a study once it reaches the recipient.

  1. Across a study, confirm that every file shows the same pseudonym for the patient.
  2. Load the study in a viewer and confirm that the series still group as they did in the original.
  3. Import the study into the target system and confirm there are no identifier collisions with studies already there.
  4. Confirm that no DICOMDIR or other index file in the release still names the patient.

Where Redactor sits in an imaging core's pipeline

Many imaging cores already run a header de-identification pipeline that keeps pseudonyms and UIDs consistent across a study, and Redactor slots into that arrangement rather than replacing any part of it. The crosswalk and any date offsets stay with that pipeline too. It hands each study to Redactor through the REST API, or places it in an Amazon S3 bucket or Azure Blob container, from which studies are ingested in bulk on a schedule. Redactor then removes the burned-in text on every frame and the PHI typed into free-text fields, applies the header rules the core has given it, and returns each file as DICOM.

Adding a redaction step to storage and systems of record that already exist is the pattern on Redactor's page for redaction in your stack.

TopicsPHI RedactionHIPAA ComplianceRedactor

You may also like

Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe

I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...

Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera

A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...

Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video

A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...

See all blogs

See it on your own content

Tell us what you are trying to solve and we will show you how it works on your infrastructure.