Benchmarking large language models for de-identification of electronic health record notes
Clinical Snapshot
PICO Framework
| P — Population | Electronic health record (EHR) clinical text notes containing sensitive health information (SHI), drawn from five benchmark datasets (i2b2-2006, MIMIC-2008, i2b2-2014, i2b2-2016, OpenDeID v1) spanning multiple countries |
| I — Intervention | Large language model (LLM)-based de-identification methods, including supervised fine-tuning approaches across eight LLM-based model configurations and nine experimental settings |
| C — Comparator | Traditional rule-based and hybrid de-identification baseline models (three baseline configurations), including single-dataset and combined-corpus training settings |
| O — Outcomes | De-identification performance measured by strict F1 micro-average score across diverse EHR datasets; cross-dataset generalisation and model robustness assessed across heterogeneous sources |
Bottom Line
This benchmarking study demonstrates that supervised fine-tuned large language models substantially outperform traditional rule-based approaches for de-identifying electronic health record notes, achieving a strict F1 score of 0.9447 compared with 0.8172 for the best baseline — a clinically meaningful 12.75 percentage point improvement. The use of combined multi-dataset training was the single most important factor driving performance in both paradigms. However, several limitations temper enthusiasm for immediate clinical deployment. The study relies on predominantly US-centric benchmark datasets, limiting generalisability to Australian or other non-US EHR environments. Precision of results is poorly characterised, with no confidence intervals reported. Critically, recall for sensitive health information — arguably the most important metric for privacy protection — is not reported separately from the composite F1 score. Performance variability across heterogeneous EHR sources remains a significant unresolved challenge. For Australian health services, local validation against Australian-specific identifiers and compliance with the Privacy Act 1988, My Health Records Act 2012, and OAIC guidelines would be mandatory before deployment. This study provides a useful technical foundation for the field but should be regarded as a research benchmark rather than a deployment-ready clinical solution.
Key Findings
P Value: Not reported
Effect Size: Best LLM-based model (supervised fine-tuning, combined corpus, setting 9): strict F1 = 0.9447; Best baseline model (combined corpus, setting 3): strict F1 = 0.8172; absolute improvement of 0.1275 F1 points (~12.75 percentage points)
Primary Outcome: De-identification performance (strict F1 micro-average score) across five benchmark EHR datasets using rule-based baseline and LLM-based models in nine experimental settings
Nnt Or Sensitivity: Sensitivity (recall) and specificity not separately reported in abstract; strict F1 score used as composite accuracy metric. For de-identification tasks, recall for sensitive health information is the critical safety parameter — not reported separately.
Confidence Interval: Not reported
Clinical Application
Fine-tuned LLM deployment requires substantial computational infrastructure, access to labelled training data, and robust data governance frameworks. The combined-corpus training approach (highest performing) requires multi-institutional data sharing for model development, which presents significant governance and consent challenges. Rule-based and hybrid approaches remain more immediately deployable in resource-constrained settings. Australian relevance is limited but growing. The My Health Record system and state-based EHR platforms (e.g., NSW HealtheNet, Victorian VHIMS) generate large volumes of clinical text requiring de-identification for secondary research use under the My Health Records Act 2012 and Privacy Act 1988. The OAIC's Australian Privacy Principles impose strict requirements on SHI handling. TGA's emerging regulatory framework for Software as a Medical Device (SaMD) may apply to clinical de-identification tools used in diagnostic or treatment contexts. RACGP standards for patient data management would require validation of any deployed de-identification system against Australian-specific identifier types (Medicare numbers, IHI, Australian addresses, state-specific facility codes). The predominantly US-centric benchmark datasets used in this study do not directly validate performance on Australian clinical text, and local validation studies would be required before implementation. The Australian Health Data De-identification Framework (AIHW) provides relevant guidance for local adaptation. Healthcare organisations seeking to de-identify EHR clinical notes for secondary use (research, quality improvement, data sharing). Most directly applicable to institutions using English-language EHR systems with structured clinical documentation similar to US benchmark datasets.
Abstract
OBJECTIVES: The rapid evolution of large language models (LLMs) and their growing application in clinical text processing have created an urgent need for reliable de-identification mechanisms. While LLMs show promise in identifying sensitive health information (SHI), their capabilities require rigorous evaluation. This study aims to conduct a comprehensive benchmarking analysis of various LLM-based, traditional rule-based and hybrid de-identification methods. METHODS: Our benchmark analysis used five datasets (i2b2-2006, MIMIC-2008, i2b2-2014, i2b2-2016 and OpenDeID v1) from different countries. We developed three baseline and eight LLM-based models. The experimental setup encompassed nine different settings using various combinations of training and testing sets to assess model robustness and cross-dataset performance. RESULTS: In the baseline models, the approach trained on the combined corpus of all five datasets (setting 3) significantly outperformed the other settings, achieving a strict F1 micro-average score of 0.8172. Regarding LLM-based models, the supervised fine-tuning approach using the same combined configuration (setting 9) achieved the highest performance with a strict F1 score of 0.9447. DISCUSSION: The harmonisation of corpora ensured standardised data formatting and SHI management across five diverse datasets, highlighting the necessity for uniform categorisation to enhance the reliability of de-identification results. CONCLUSIONS: Our findings indicate that while fine-tuned LLMs offer superior accuracy, the observed performance variability across heterogeneous electronic health record sources poses significant technical challenges. Real-world implementation must address these inconsistencies to overcome the ethical and technical hurdles associated with deploying LLMs for handling sensitive health data.
References
- 1.Panchal, O., Chang, N.-W., Zhao, Z.-R., Nadar, D. R., Dai, H.-J., & Jonnagaddala, J. (2026). Benchmarking large language models for de-identification of electronic health record notes. BMJ Health & Care Informatics. https://doi.org/10.1136/bmjhci-2025-101894
Related Research
Clinical gastroenterology and hepatology : the official clinical practice journal of the American Gastroenterological Association
Artificial Intelligence Tools for Gastrointestinal Research: A Practical Guide
2 Aug 2026
Clinical spine surgery
An Introduction to Machine Learning for the Practicing Spine Surgeon
2 Aug 2026
Rheumatic diseases clinics of North America
Sources of Bias in Clinical Artificial Intelligence and Applications in Rheumatology
2 Aug 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service