PHI Redaction, HIPAA Compliance, Redactor
Which DICOM Tags Contain PHI, and What to Replace Them With
If your imaging core de-identifies studies with its own script, the heart of that script is a list of DICOM (Digital Imaging and Communications in Medicine) tags. Two questions decide whether that list does its job, namely which DICOM tags contain protected health information (PHI) and what to write in their place once a value has to go. A cleaned header can still be a broken one, since a date field holding the word REDACTED can stop strict viewers and picture archiving and communication systems (PACS) from reading the file.
The tags below are grouped by what they identify and mapped to the Safe Harbor categories of the Health Insurance Portability and Accountability Act (HIPAA). Our guide to DICOM de-identification sets those categories out alongside the pixels and the rest of the file.
The DICOM tags that carry PHI, grouped by what they identify
Every attribute in a DICOM header has a tag made of a group number and an element number, such as (0010,0020) for Patient ID, and most patient attributes sit in group 0010. The complete reference is Table E.1-1 in PS3.15, which lists every attribute the standard's confidentiality profile acts on. The table below groups the attributes a Safe Harbor review meets most often. Table E.1-1 lists many more, among them Patient's Birth Time (0010,0032), Other Patient IDs (0010,1000) and the equipment's Device Serial Number (0018,1000). A script should be checked against the full table, not against a short list, this one included.
| What it identifies | Attributes and tags | Safe Harbor category |
|---|---|---|
| The patient by name | Patient's Name (0010,0010), Other Patient Names (0010,1001) | (A) Names |
| The patient's record | Patient ID (0010,0020), Other Patient IDs Sequence (0010,1002), Medical Record Locator (0010,1090) | (H) Medical record numbers |
| Who assigned the ID | Issuer of Patient ID (0010,0021) | Not listed, but it names the source of the record |
| Birth and age | Patient's Birth Date (0010,0030), Patient's Age (0010,1010) | (C) Dates, and ages over 89 |
| Where the patient lives | Patient's Address (0010,1040), Region of Residence (0010,2152) | (B) Anything smaller than a state |
| The patient's country | Country of Residence (0010,2150) | Not listed, since a country is larger than a state |
| How to reach the patient | Patient's Telephone Numbers (0010,2154), Patient's Telecom Information (0010,2155) | (D) Telephone numbers, (F) Email addresses |
| Other demographics | Occupation (0010,2180), Ethnic Group (0010,2160), Patient's Religious Preference (0010,21F0) | Not listed |
| Physicians and operators | Referring Physician's Name (0008,0090), Performing Physician's Name (0008,1050), Name of Physician(s) Reading Study (0008,1060), Physician(s) of Record (0008,1048), Operators' Name (0008,1070) | Not listed |
| The institution | Institution Name (0008,0080), Institution Address (0008,0081) | Not listed |
| When the imaging happened | Study Date (0008,0020), Series Date (0008,0021), Acquisition Date (0008,0022), Content Date (0008,0023) | (C) Dates |
| The order and the exam | Accession Number (0008,0050), Study ID (0020,0010) | (R) Other unique identifying numbers |
Several rows carry no Safe Harbor letter, and they still belong on the list. The DICOM profile removes physicians, operators and the institution anyway, as our guide explains, because together they can point to one patient, and it removes Country of Residence as well. An unusual occupation or a rare ethnic group in a small population can also narrow a study to very few people. A de-identification program has to decide how it treats these fields instead of keeping them by default.
Free text is where a tag list runs out
Free-text attributes hold whatever someone typed, which makes them the weakest point of any list that works by field name. Study Description (0008,1030) and Series Description (0008,103E) are meant to describe the exam, but nothing stops a technologist typing a name or a date of birth into them. Image Comments (0020,4000), Patient Comments (0010,4000) and Additional Patient History (0010,21B0) are open text by design.
A small test study shows the gap in any tool that works by field name alone. Put a name and a date of birth in the standard patient fields, and type the same details into the study description, the series description and the image comments. Then run the tool with its rule list alone. The standard fields come back masked while the three free-text fields come through exactly as typed, because nothing in a list of field names tells the tool to read what someone wrote. Run it again with detection reading the text the rules do not name, and the typed copies should be gone as well.
The standard's baseline takes the blunt route with these fields, and Table E.1-1 marks Image Comments for removal. Its Clean Descriptors Option exists for projects that need descriptive text kept, and it cleans the text instead of removing it. Detection is what makes that cleaning practical across thousands of studies, and our article on how named entity recognition finds PHI in clinical text explains how such models find names and dates inside sentences.
Remove, empty or replace: what the standard allows
Once you know which attributes identify someone, the standard gives each one an action code, and the right choice depends on whether the attribute has to stay in the object. The definitions below are quoted from PS3.15, where VR stands for value representation, the attribute's data type.
| Code | Definition in PS3.15 | Examples under the Basic Profile |
|---|---|---|
| X | "remove Attribute", including every item of a sequence | Other Patient IDs Sequence, Patient's Address, Image Comments |
| Z | "replace with a zero length value, or a non-zero length value that may be a dummy value and consistent with the VR" | Patient's Name, Patient's Birth Date, Accession Number |
| D | "replace with a non-zero length value that may be a dummy value and consistent with the VR" | Content Sequence in structured reports |
| C | "clean, that is replace with values of similar meaning known not to contain identifying information and consistent with the VR" | Image Comments, when the Clean Descriptors Option applies |
| K | "keep", unchanged for ordinary attributes and cleaned for sequences | Attributes that a chosen option retains |
| U | "replace with a non-zero length UID that is internally consistent within a set of Instances" | Unique identifiers that group and reference instances |
Deleting an attribute is not always allowed, because PS3.5 makes some of them mandatory in the objects that carry them. Type 2 attributes "shall be included in the Data Set and their absence is a protocol violation", although they may be sent empty when the value is unknown. Type 1 attributes must carry a value, since "The Length of the Value Field shall not be zero." Patient's Name, Patient ID, Patient's Birth Date and Accession Number are all Type 2 in the modules that define them, which fits the Z the baseline gives most of them. For protected Type 1 attributes the standard adds that "Dummy values may be necessary", because an empty value is not allowed.
Replacing the value while leaving the attribute in place avoids most of these traps, since software that expects an attribute keeps finding it. Putting a detector on the header raises a different risk, because a header is full of short codes and numbers that a detector trained on prose can mistake for identifiers. Modality, Photometric Interpretation, Number of Frames and instance numbers are examples, and masking one of them leaves a file that no longer describes its own pixels. Only text fields that can carry a person's details should ever reach a detector.
A replacement value has to fit the field's data type
Every DICOM attribute declares a value representation, and PS3.5 fixes the format each one may hold, which is the rule any replacement value has to respect.
| VR | What the standard allows | Replacements that break it |
|---|---|---|
| DA, date | "A string of characters of the format YYYYMMDD", 8 bytes fixed | The word REDACTED, or a date written with slashes |
| TM, time | "HHMMSS.FFFFFF", with hours from 00 to 23 | Letters in place of digits |
| DT, date-time | "YYYYMMDDHHMMSS.FFFFFF&ZZXX", at most 26 bytes | A date in day-month-year order |
| AS, age | "nnnD, nnnW, nnnM, nnnY", 4 bytes fixed, so 018M means 18 months | "90+", "over 89", or a number with no unit |
| PN, person name | Family, given, middle, prefix and suffix components separated by ^, at most 64 characters per group | Text longer than the limit |
| CS, code string | Uppercase letters, digits, space and underscore, at most 16 bytes | Lowercase text or punctuation, such as n/a |
Every replacement therefore has to suit the field's data type, since a text mask in a date or age field leaves a file that strict parsers reject even with every identifier gone. A person-name field can take placeholder text, while a date field can only take something shaped like a date, so no single mask can serve both.
Dates need a decision before any value is written, because Safe Harbor's date rule, covered in our guide, keeps only the year. A fixed neutral date in the right format removes the date, and it also removes the interval between a patient's scans, which a longitudinal study may need. Keeping those intervals is a separate problem that belongs with the pipeline holding the patient crosswalk, and the standard's option for it is set out in our article on multi-frame files, series and studies.
Ages over 89 are awkward in DICOM, because the age format holds only a three-digit number and a unit. A program that keeps ages can write one agreed value, such as 090Y, for every patient over 89 and record what it means, while the baseline profile simply removes Patient's Age.
How to check a de-identified header
A header check has to prove two things at once, that every identifying value is gone and that everything else still reads as valid DICOM. Each attribute in the table of PHI tags above should be either absent or holding a value its data type allows. Structural attributes such as Modality, Photometric Interpretation, Rows, Columns and Pixel Spacing should be exactly as they were. A DICOM validator does much of this work in one run, reporting missing Type 2 attributes, empty Type 1 attributes and values in the wrong format before a receiving PACS rejects the file.
A header can also record that it was de-identified, and a release process should decide who sets that record and how it is checked. Under the baseline profile, Patient Identity Removed (0012,0062) "shall be replaced or added to the Data Set with a value of YES". The profile also asks for the method to be recorded, as codes in De-identification Method Code Sequence (0012,0064), as a text description in De-identification Method (0012,0063), or both. A recipient can read that value without opening anything else, so it should be set by the step that actually did the work and checked like any other output. Copies hiding in private groups and nested items call for the separate checks in our article on DICOM private tags and nested sequences.
What Redactor changes in a DICOM header
VIDIZMO Redactor handles a DICOM header in two layers, a standard rule list for the fields known to identify someone and an optional artificial intelligence (AI) pass over the text the list does not name.
| Header content | What Redactor does with it |
|---|---|
| Fields on the rule list, covering the patient, physicians and operators, the institution, dates and times, and record and accession numbers | Keeps the tag and replaces its value with one legal for the field's data type |
| Free text the list does not name, such as study and series descriptions and image comments | Reads it in the optional AI pass, and a field flagged there has its value replaced the same way |
| Binary elements, and text fields that cannot hold a person's details | Never offers them to detection |
The same Safe Harbor identifiers in records, recordings and documents are covered on Redactor's HIPAA redaction page.
TopicsPHI RedactionHIPAA ComplianceRedactor
About the author
Nabeel Ali is a Senior Product Analyst at VIDIZMO, with five years at the company. An engineer by profession, Nabeel works on audio, video and DICOM redaction in Redactor, and has also worked on AI Intelligence Hub and AI Live Insight. The role sits where customers and engineering meet: hearing from customers, customer success, and the sales and marketing teams what goes wrong in real deployments, then working with the engineers until the product answers it.
You may also like
Video Redaction Best Practices: Motion, Frame Rates, Tracking and Failing Safe
I led our video redaction project for a county public safety agency where two people handled every disclosure, and ...
Redacting Dash Cam, Body Cam and Drone Footage From a Moving Camera
A fleet claims manager preparing crash footage for an insurer and defense counsel is doing a different job from a ...
Why Frame-Rate Headers Lie: Redacting Variable Frame Rate Video
A video file keeps time frame by frame, and the frames-per-second figure a player displays is a summary of that timing, ...
See it on your own content
Tell us what you are trying to solve and we will show you how it works on your infrastructure.