How Medical Imaging De-Identification Powers AI Research and HIPAA Compliance

De-Identification In Medical Imaging

De-identification removes or transforms personally identifiable information in medical images so the data can be used for AI development, clinical research, quality improvement, and other authorized secondary purposes while protecting patient privacy. Unlike conventional data files, medical images may contain identifiers in both DICOM metadata and text embedded within the image pixels.

An effective imaging workflow must remove or transform direct identifiers defined under HIPAA, including names, medical record numbers, contact information, and relevant dates, while preserving the clinical content needed for analysis. This requires validation across metadata, pixel data, and related files rather than relying on a single masking step.

Large initiatives and benchmarking challenges demonstrate the need for scalable automation combined with quality assurance and, when necessary, human review. The HIPAA Privacy Rule allows covered entities and business associates to de-identify health information according to defined standards, enabling properly prepared imaging datasets to support research and innovation without exposing individual patients.

Medical image before de-identification. Source: Dicom Systems

De-Identification Techniques

De-identification methods fall into one of two categories, depending on where the PHI is stored within the file. In medical imaging, patient data can be embedded in the image (also known as burned-in annotations) or stored as metadata tags in file properties. Identification of patient and date as text in an encapsulated document (e.g., in an XML attribute or element) is equivalent to “burned in” annotation.

Chest x-ray with burned-in annotation. Source: Research Gate

Each DICOM file contains extensive metadata in a header. Metadata is information that describes the image. The header portion of a DICOM file typically contains PHI. Metadata is a powerful tool for annotating and leveraging image-related information for clinical and research purposes, as well as for organizing and retrieving archived images and their associated data.

DICOM object with metadata and image. Source: Research Gate

Historically, de-identification methodologies relied on image masking to remove burned-in data. Image masking can be a manual process that must be completed on each image. Based on the coordinates for a particular equipment model, an editor can locate the identifiers in the file and mask them. Methods of de-identification include blurring, pixelating, or blocking.

Data masking has several limitations, primarily due to the effort required and the potential for errors or missed fields. To replace data masking, automated tools that use computer vision and optical character recognition (OCR) have been developed. These de-identification tools can distinguish between patient information and clinically relevant data, automatically masking the former while preserving the latter, all in a fraction of the time it takes a human to perform the same task.

How to De-identify Data

Data de-identification is typically managed in a two-step process.

The first step in data de-identification is to classify and tag direct and indirect identifiers. Identifiers that are unique to a single individual, such as Social Security numbers, passport numbers, and taxpayer identification numbers, are known as “direct identifiers.” The remaining types of identifiers are known as “indirect identifiers” and generally consist of personal attributes that are not unique to any individual. Examples of indirect identifiers include height, ethnicity, hair color, and more. However, although not exceptional, indirect identifiers can be combined to identify an individual’s records.

Once the data classifiers have been verified and accurately represent the data sources, the tagging process can be automated. This makes the de-identification process vastly more efficient for data teams.

Data can then be de-identified through the combination of various dynamic data masking techniques and data access controls. These technical and organizational measures affect both the data’s appearance and its environment, including who can access the data and for what purposes, among other factors.

A two-stage de-identification process for CT scan images is available in DICOM file format. In the first stage of the de-identification process, the patient’s PHI, including name, date of birth, and other identifying information, is removed at the hospital facility using the export process available in its Picture Archiving and Communication System (PACS). The second stage employs the proposed DICOM de-identification tool for an exhaustive attribute-level investigation to further de-identify and ensure that all PII has been removed. Finally, we provide a roadmap for future considerations to develop a semi-automated or automated tool for de-identifying DICOM datasets.

Medical image after de-identification. Source: Dicom Systems

Medical Image De-Identification in the Cloud

The adoption of cloud computing in healthcare is enabling greater integration and collaboration between hospitals, medical organizations, and healthcare providers. The benefits of storing patient data in the cloud include enhanced security, scalability, reduced IT costs, greater availability and reliability, and 24/7 access from multiple locations.

Regarding the de-identification of data stored in the cloud, adoption is strongly linked to compliance with the Safe Harbor framework. Medical image de-identification is gaining adoption among healthcare organizations that store data in a private cloud. A private cloud is essentially an extension of the client’s data center and, therefore, is governed and protected by the same security and privacy methods as on-premises storage. In contrast, in a public cloud such as Google Cloud, Amazon Web Services, or Microsoft Azure, data de-identification would violate HIPAA’s Safe Harbor.

Re-Identification

Covered entities may also wish to re-identify data at a later date. Data re-identification, also known as de-anonymization, is the practice of matching anonymous data (also referred to as de-identified data) with publicly available information or auxiliary data to discover the individual to whom the data belongs.

If a covered entity or business associate successfully undertook an effort to identify the subject of de-identified information it maintained, the health information would again be protected under the Privacy Rule because it would meet the definition of PHI. Disclosure of a code or other means of record identification designed to enable coded or otherwise de-identified information to be re-identified is also considered a disclosure of PHI.

To prepare for re-identification, covered entities may assign a unique code to a dataset or specific records, provided that the code is not derived from information about the individual and the entity does not disclose the code or the mechanism for re-identification to any third party.
If the data is re-identified, it is considered PHI again under HIPAA.

Results Achieved with Dicom Systems De-Identification Framework for HSS Global Innovation Institute

Hospital for Special Surgery (HSS) partnered with Dicom Systems to route and de-identify over 5.3 million orthopedic imaging exams to train AI models for fracture detection. Dicom Systems deployed a customized, enterprise-grade de-identification framework on its Unifier® platform. The architecture supported scalable ingestion, metadata scrubbing, pixel data redaction, and compliant export, with strict QA enforcement at every step.

Dicom Systems delivered transformative results for the HSS Global Innovation Institute, successfully processing 5.3 million medical imaging exams while fully complying with HIPAA Safe Harbor standards. This massive-scale effort, validated by third-party auditors, enabled the secure provision of precise clinical datasets to AI partners developing algorithms to combat fracture misdiagnosis at the world’s leading musculoskeletal health center.

Results achieved include:

  • 5.3M images de-identified
  • Successful validation by 3rd third-party auditing agency
  • AI Algorithms are safely provided with the precise type and volume of clinical data required for effective machine learning

De-Identification As an On-Ramp to AI and Machine Learning Innovation

Artificial intelligence (AI) is quickly becoming an integral part of modern healthcare. AI algorithms and other applications are being used to support medical professionals in clinical settings and in ongoing research. In medical imaging, AI tools are being used to analyze CT scans, X-rays, MRIs, and other images for lesions or other findings that a human radiologist might miss. Artificial intelligence can enhance medical imaging for screenings, precision medicine, and risk assessment.

High-quality data is a crucial contribution to improving machine learning algorithms. The greater the volume and variety of data available to an algorithm, the more accurate its learning and the more precise its outcomes will be. For AI algorithms to properly “consume” or “ingest” patient data, the data must be standardized and HIPAA-compliant. To that end, Unifier’s de-identification functionality delivers de-identified data in the appropriate format and quality for use in any relevant AI algorithm.

Dicom Systems Unifier platform AI Conductor

Dicom Systems Unifier De-Identification Benefits and Features

The Dicom Systems Unifier platform supports de-identification of a wide range of image formats, including DICOM, XML, TIFF, JPEG, and PDF, ensuring compliance with HIPAA Safe Harbor requirements for Protected Health Information (PHI). The platform processes incoming images and data, translating them into a standardized format that can be securely accessed or transferred to referring physicians, radiologists, PACS/MIMPS, RIS, or any radiology workstation, regardless of their physical location. This flexibility, combined with robust automation, allows healthcare organizations to protect patient privacy while maintaining interoperability and clinical workflow efficiency across diverse systems.

Dicom Systems Unifier Offers Deidentification Features and Benefits:

  • Adherence to the HIPAA Privacy Rule and Safe Harbor requirements by third parties.
  • Full customization of processes and output
  • Robust enough for large-scale de-identification with no impact on clinical workflow.
  • Supports full DICOM, DICOMweb, FHIR, and HL7 interoperability with compatible devices
  • Removes pixel-, mask-, metadata-, and text-based information from images, referencing an industry database of modalities by vendor and model to identify the exact coordinates at which text was burned into medical images. Optical Character Recognition (OCR) software capabilities are available for specific use cases.
  • Best price-to-performance technology trusted by top healthcare enterprises, government agencies, and imaging partners.
  • When deployed with Dicom Systems Unifier Vendor-Neutral Archive (VNA), it leverages a robust framework for imaging lifecycle management and archiving.
  • Bidirectional dynamic tag morphing affects both input and output.
  • Advanced pixel-level de-identification to avoid accidental corruption or truncation of the image file.
  • Complex DICOM tag substitutions, removals, or morphing are automated by defining transformations within the Lua script framework.
  • Full customization of de-identification processes and output.
De-identification unlocks the full potential of healthcare data by safeguarding patient privacy while enabling its reuse in groundbreaking research, AI development, and public health initiatives. Partner with Dicom Systems to securely de-identify your medical imaging data at scale, ensuring compliance with privacy regulations while accelerating innovation and collaboration in healthcare.
Meet with one of our enterprise imaging workflow experts.