Evidence-Based Medicine

Research Appraisals

Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.

Showing 1 appraisal

otherEvidence: Moderate
60CEBM

BMJ health & care informatics

Benchmarking large language models for de-identification of electronic health record notes

OBJECTIVES: The rapid evolution of large language models (LLMs) and their growing application in clinical text processing have created an urgent need for reliable de-identification mechanisms. While LLMs show promise in identifying sensitive health information (SHI), their capabilities require rigorous evaluation. This study aims to conduct a comprehensive benchmarking analysis of various LLM-based, traditional rule-based and hybrid de-identification methods. METHODS: Our benchmark analysis used five datasets (i2b2-2006, MIMIC-2008, i2b2-2014, i2b2-2016 and OpenDeID v1) from different countries. We developed three baseline and eight LLM-based models. The experimental setup encompassed nine different settings using various combinations of training and testing sets to assess model robustness and cross-dataset performance. RESULTS: In the baseline models, the approach trained on the combined corpus of all five datasets (setting 3) significantly outperformed the other settings, achieving a strict F1 micro-average score of 0.8172. Regarding LLM-based models, the supervised fine-tuning approach using the same combined configuration (setting 9) achieved the highest performance with a strict F1 score of 0.9447. DISCUSSION: The harmonisation of corpora ensured standardised data formatting and SHI management across five diverse datasets, highlighting the necessity for uniform categorisation to enhance the reliability of de-identification results. CONCLUSIONS: Our findings indicate that while fine-tuned LLMs offer superior accuracy, the observed performance variability across heterogeneous electronic health record sources poses significant technical challenges. Real-world implementation must address these inconsistencies to overcome the ethical and technical hurdles associated with deploying LLMs for handling sensitive health data.

23 July 2026

Read appraisal →