PHI Redaction, HIPAA Compliance, Redactor

DICOM De-identification: Tags, Burned-In Text and the Pixels In Between

My team built DICOM de-identification into VIDIZMO Redactor because the organizations asking us for it kept running into the same gap. Clinical research organizations running multi-site trials, a specialty care center contributing scans to multi-center studies, and teams assembling imaging to train AI models all needed identifiers gone from the header and from the text burned into the pixels. That had to hold on every frame of every study before the imaging left their hands, and the header scrubbers they already ran left the burned-in text where it was.

Research isn't the only reason imaging leaves a hospital, and not every release is de-identified. Second opinions and transfers keep the patient's identity, because the receiving team has to know whose scan it is, and releases to law firms and insurers carry only what the request or the patient's authorization covers. For research itself, de-identification is one of several routes, and a patient's authorization, a waiver from an institutional review board or privacy board, and a limited data set under a data use agreement are others.

In my experience, few redaction vendors handle DICOM at all, and many of those that do clean the header and leave the text burned into the pixels. The pixels are the hard part, because a study holds many slices and frames and the burned-in text changes from one to the next. The text therefore has to be found on every frame rather than copied forward from the first. Pixel tools built on templates blank a fixed region for each scanner model, which works for a banner that never moves and misses text anywhere else.

Where PHI hides in a DICOM file

A DICOM file keeps protected health information (PHI) in three places at once, and de-identification has to reach all three whichever kind of release it serves.

  • The header. Named attributes such as Patient's Name, Patient ID, birth date, study dates, referring physician and accession number, and free-text fields like Study Description that hold whatever someone typed at the scanner.
  • The pixels. Text burned into the image, routine on ultrasound and common on screenshots and scanned forms saved as DICOM, and in head scans a face that can be rebuilt from the image data.
  • The structures in between. Sequences nested inside other sequences, private attributes that each manufacturer defines for itself, and overlays, presentation states and structured reports that travel with the image.

Header work starts from a list of tags known to identify someone, and that list covers only the fields someone thought to name. Free text runs past any list, since no field name predicts what a technologist will type, so the text itself has to be read. A removed value also needs a replacement that fits the field's data type, or strict viewers and picture archiving and communication systems (PACS) can refuse a file whose identifiers are all gone. Nabeel Ali on my team, who worked on our DICOM redaction, has mapped the DICOM tags that contain PHI, with a valid replacement for each data type.

A header script that reads only the top level of the header lets identifiers through in two places. A manufacturer's private field can hold a copy of the patient's name, and the same name can sit again several levels down inside a sequence. Some of what those structures carry is exactly what the image needs, like the ultrasound calibration that lets a viewer measure, so stripping every sequence to be safe breaks the study for its recipient. Private tags and nested sequences need a policy of their own for exactly that reason.

What HIPAA and the DICOM standard ask for

HIPAA, the Health Insurance Portability and Accountability Act, offers two methods of de-identification, and neither says anything about DICOM or any other file format. Safe Harbor removes 18 kinds of identifier, lettered (A) to (R) in 45 CFR 164.514(b)(2), for the patient and for relatives, employers and household members. It also requires that the organization has no actual knowledge that what remains could identify the person. Expert Determination relies instead on a qualified expert finding a very small risk that an anticipated recipient could identify someone, which leaves room to keep detail a study needs.

Safe Harbor names kinds of information rather than places in a record, so a name burned into an ultrasound frame counts as much as the Patient's Name attribute. The table maps the identifiers that matter most for imaging to the places they turn up.

Safe Harbor identifier In the header In the pixels and elsewhere
(A) Names Patient's Name, Other Patient Names, names typed into descriptions and comments Ultrasound banners, worklist screenshots, scanned forms
(B) Geographic units smaller than a state Patient's Address, Region of Residence Scanned referral letters and forms
(C) Dates other than the year, and ages over 89 Birth date, age, and the study, series, acquisition and content dates Date and time stamps burned into frames
(D) to (F) Telephone, fax and email Patient's Telephone Numbers, Patient's Telecom Information Scanned documents saved as DICOM
(G), (I) and (J) Social Security, health plan and account numbers Free text, wherever someone typed them Scanned requisitions and insurance forms
(H) Medical record numbers Patient ID, Other Patient IDs Sequence, Medical Record Locator ID lines on banners and screen captures
(Q) Full-face photographs and comparable images None, since this is image content Clinical photographs, and faces rebuilt from head scans
(R) Any other unique identifying number or code Accession Number, Study ID Burned-in text on screenshots and forms

Dates need the most care, because one study carries many of them and Safe Harbor keeps only the year of any date directly related to the patient. Ages over 89 go too, along with every part of a date that shows such an age, year included, unless they're grouped into a single category of 90 or older.

The DICOM standard answers in PS3.15 Annex E, whose Basic Application Level Confidentiality Profile describes itself as "an extremely conservative approach". Besides the patient's identity, it removes the identity of the personnel involved in a procedure and of the organizations that ordered or performed it. Safe Harbor leaves those out because they aren't the patient's own identifiers, yet together they can narrow a study down to one person.

One sentence in the profile explains why header-only de-identification is both common and risky. It reads, "Unless the Clean Pixel Data Option or the Clean Recognizable Visual Features Option is specified, this Profile does not address information in the pixels." A tool can follow the baseline to the letter and still ship a name burned into the image, since cleaning the pixels is an option a project has to choose. Even the header is harder than it looks, and in a study published in 2015, Aryanto and colleagues found only one of ten free toolkits removing all fifty identifying header elements at its default settings.

Many tools call the job DICOM anonymization, and the two words describe the same work in practice, although the standard and HIPAA both say de-identification. A release is judged by the standard its output meets, whether that's Safe Harbor, Expert Determination or a named DICOM profile. We don't claim conformance to a named DICOM profile, which is one reason research cores run Redactor inside their own pipeline.

What we built, and why it runs in one pass

Because identifiers sit in the header and the pixels at the same time, my team treats the two surfaces as one job over the native file. A multi-frame file can run to several hundred megabytes, and two separate tools would rewrite it twice, while a conversion to images or PDF hands the recipient something that is no longer DICOM. A failed second pass could also leave a file half-processed, so Redactor opens each file once, de-identifies both surfaces and saves it once as valid DICOM.

In the header, a rule list covers the fields known to identify someone and reaches every copy of a field, including copies nested inside sequences. Each field it names stays in place as a tag and gets a replacement value its data type allows. An optional AI pass reads the free-text fields the list doesn't name, such as study descriptions and image comments, and a field it flags has its value replaced the same way.

In the pixels, Redactor reads every frame for burned-in text, including text that isn't upright, and masks only what detection classifies as personal, so orientation markers and measurements stay readable. The mask is written into the stored values in the image's own bit depth and signedness, because a graphic drawn over the text would leave the original values in the file. Redacting DICOM pixel data without breaking the image goes through the values involved and the comparisons that prove a mask worked.

After masking, the pixel data is written uncompressed, which keeps every unmasked value exactly as it was decoded from the original file. A compressed study therefore comes back larger, so I'd plan storage and transfer for that before the first large release.

Identified releases: keeping the patient, removing everyone else

A second opinion, a transfer or an attorney's request goes out identified, and a configuration built for research would strip the very name the release exists to carry. Which header tags get replaced is decided by a rule list, and a portal's administrator sets it once for everything that portal processes. So I'd run identified releases from a portal of their own, whose rules leave the header's own fields in place: the patient's name, ID and dates, and the staff and institution who produced the study. A transfer or a subpoenaed copy has to be the true record. Research and teaching releases stay in another portal with the full rules, where private tags are removed unless they're named on a keep list.

The pixels and any tag the rules don't name go through detection, and its settings carry a list of words it skips. A name on that list is never flagged, so the identified-release portal's list can carry the department's own staff permanently, plus the patient whose release is running. Every other name detection finds is still masked, such as another patient's on a worklist screenshot. The list sits in the processing settings rather than on each request, so that portal runs one identified release at a time. The patient's name comes off the list when a release goes out, because a name left over from an earlier release would be skipped in the next one.

A person should still look at every image before an identified release goes, because nothing in a DICOM file records whose details a screenshot or a scanned page holds. Other patients' details tend to arrive as burned-in text in DICOM secondary captures, carried in by worklist screenshots, dose screens and scanned forms filed into the wrong study. Legal and insurer releases add a judgment no tool makes, since someone has to read the request or the authorization and decide what it covers before the configuration is set.

Research releases, beside the honest broker's pipeline

In many imaging cores, de-identification for research already belongs to an honest broker, whose crosswalk maps each patient to a study pseudonym and stays out of the research team's reach. A header pipeline applies that crosswalk, the replacement unique identifiers (UIDs) and any date shift across every file of a study. For a core like that, I'd put Redactor beside the pipeline rather than in place of it, and the crosswalk, pseudonyms, UIDs and any date shifting stay with the organization. Redactor doesn't remove faces that can be rebuilt from head scans, so defacing stays a step of its own in that pipeline.

Redactor takes on the work a header pipeline was never built for, removing burned-in text from every frame and PHI typed into free-text fields. It applies whatever header rules the core gives it, so fields the pipeline already handles can be left to the pipeline. Studies arrive through the REST API, or they're pulled in bulk from Amazon S3 or Azure Blob storage on a schedule, and each file goes back as DICOM.

Across a whole study, the de-identified pieces still have to fit together when they arrive. Every file needs the same pseudonym for the same patient, replacement UIDs have to keep series grouped without colliding in the receiving PACS, and date shifts have to be applied consistently for each patient. A patient split across two pseudonyms can land in both the training set and the test set of a model, and a DICOMDIR index file can still name the patient after every image is clean. Multi-frame files, series and whole studies each call for checks of their own, from the frames inside one file to a study spread across hundreds of files.

How to check DICOM de-identification before a study leaves

I'd judge every tool, ours included, by what its output contains, and the checks below need nothing more than a DICOM toolkit, a viewer and the system the files are going to.

  1. Dump every element of the output, including the items inside sequences and every private group, and read the dump instead of a viewer's header panel.
  2. Search the raw bytes of each file for the patient's name, record number and date of birth, which catches copies a header panel never shows.
  3. Step through every frame at several window settings, looking for names, numbers and dates, since text in a dark corner can disappear at one setting and reappear at another.
  4. Open the file in a strict viewer and import it into the receiving system before the whole release goes.

Run the checks on a sample from every modality and every contributing site, because each scanner model and export path has habits of its own. An identified release adds one question for the reviewer, which is whether every name and record number visible on the images belongs to the patient on the request. Detection also depends on how your devices print their text, so we work with each organization to calibrate it to its own studies, and a pilot on them measures the result.

What to license, and where it runs

For everything described here you license VIDIZMO Redactor, which runs on the VIDIZMO platform that stores the studies, controls who can open them and connects to the storage they come from. It can be hosted by VIDIZMO, run in your own cloud subscription or data center, or run air-gapped with every AI model local, and no other product is needed for DICOM. VIDIZMO provides data-protection, security, redaction, and anonymization capabilities for PHI/PII that help customers meet HIPAA obligations, and can enter into a Business Associate Agreement (BAA) where required.

How Redactor treats the pixels and the tags of a study is summarized on its DICOM redaction page.

TopicsPHI RedactionHIPAA ComplianceRedactor

You may also like

Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe

I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...

Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera

A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...

Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video

A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...

See all blogs

See it on your own content

Tell us what you are trying to solve and we will show you how it works on your infrastructure.