Research Appraisalother

Natural language processing tool for extracting information about opioid overdoses in the USA from case narratives in the violent death reporting system.

Injury prevention : journal of the International Society for Child and Adolescent Injury PreventionOmaki, Elise, Restrepo, Felipe, Shields, Wendy C et al.20 July 2026DOI

Clinical Snapshot

65CEBM
Evidence: Moderateother

PICO Framework

P — PopulationDeath records of drug overdose victims in the USA, sourced from the National Violent Death Reporting System (NVDRS)
I — InterventionA rule-based natural language processing (NLP) tool ('Excel Extractor') using spreadsheet formulae to extract and code 12 attributes across 82 labels from free-text case narratives
C — ComparatorManually labelled (human-annotated) data as the reference standard; multiple machine learning (ML) models including deep learning architectures
O — OutcomesAccuracy of information extraction measured by F1 Score across 12 attributes and 82 labels; comparative performance against ML models

Bottom Line

This methodological study from Johns Hopkins and Virginia Tech demonstrates that a rule-based NLP tool built using spreadsheet formulae — the 'Excel Extractor' — can accurately extract structured information about drug overdose circumstances from free-text death record narratives in the US National Violent Death Reporting System. Across 12 attributes of interest, the tool achieved F1 Scores above 0.8 on nine, and outperformed multiple machine learning models on seven. Notably, 30% of individual labels achieved F1 ≥ 0.9. The key practical advantage is accessibility: the tool requires no specialised computational infrastructure, making it deployable in resource-limited public health agencies. However, the study has important methodological gaps. Inter-rater reliability of the manual reference standard is unreported, confidence intervals around performance estimates are absent, and external validation has not been performed. Three attributes underperformed the 0.8 threshold, and the characteristics of these failures are not described. For Australian practitioners, the tool is not directly transferable to the National Coronial Information System without local re-validation, but the methodological approach is instructive for improving opioid overdose surveillance timeliness. This is a promising proof-of-concept warranting prospective validation before operational deployment.

Evidence: Moderate

Key Findings

  • P Value: Not reported

  • Effect Size: F1 Score > 0.8 achieved on 9 of 12 attributes; F1 ≥ 0.8 on 46 of 82 labels (56%); F1 ≥ 0.9 on 25 of 82 labels (30.5%); highest performing model on 7 of 12 attributes compared to ML benchmarks

  • Primary Outcome: Accuracy of the Excel Extractor in extracting 12 attributes (82 labels) about drug overdose circumstances from NVDRS free-text narratives, measured by F1 Score

  • Nnt Or Sensitivity: F1 Score (harmonic mean of precision and recall) used as primary performance metric; individual precision and recall values not reported in the abstract. F1 > 0.8 is generally considered acceptable for clinical NLP information extraction tasks.

  • Confidence Interval: Not reported

Clinical Application

High feasibility for adoption in settings with existing Microsoft Excel infrastructure and without access to advanced computational resources. The rule-based approach requires domain expertise to develop and maintain rule sets but does not require machine learning expertise or specialised software. Maintenance burden as narrative styles evolve over time is a consideration. Australia's primary equivalent infrastructure is the National Coronial Information System (NCIS), which contains free-text coronial narratives for drug-related deaths. The Australian Institute of Health and Welfare (AIHW) and state-based coroners' courts manage these data. A similar rule-based NLP approach could theoretically be adapted for NCIS narratives to improve timeliness and granularity of opioid overdose surveillance, which is a recognised priority given Australia's evolving illicit drug landscape including fentanyl analogues and pharmaceutical opioid misuse. The TGA's pharmacovigilance activities and the RACGP's opioid prescribing guidelines both depend on timely surveillance data. However, direct application of the US-trained Excel Extractor to Australian data would require substantial re-labelling and re-validation given differences in coronial narrative structure, terminology (e.g., 'Schedule 8' vs. DEA scheduling), and reporting conventions. PBS opioid prescribing data linkage with NCIS narratives could further enhance the utility of such a tool in the Australian context. Public health epidemiologists, medical examiners, coroners, and injury surveillance professionals working with free-text death record narratives in drug overdose surveillance systems

Abstract

BACKGROUND: Improving the infrastructure for drug overdose surveillance is critical for identifying new threats and responding to emerging trends. We aimed to develop a prototype tool using the principles of natural language processing that can extract information from the death records of drug overdose victims. METHODS: Data were obtained from the Violent Death Reporting System on drug overdose deaths. Narratives were manually labelled for 12 attributes of interest, totalling 82 labels about the circumstances of the overdose. Narratives were passed through the 'Excel Extractor' to identify and extract a target phrase and subsequently map the extracted phrase to predetermined code values. The output from the Excel Extractor was compared with manually labelled data to determine accuracy. Performance was compared against multiple machine learning models. RESULTS: The Excel Extractor performed well across the attributes of interest, achieving an F1 Score over 0.8 on nine of the 12 attributes. The Excel Extractor was the highest performing model on seven of the 12 attributes. The Excel Extractor achieved an F1 Score of 0.8 or higher on 46 of 82 (56%) of the labels, and a score of 0.9 or higher on nearly one-third (25 out of 82) of the labels. CONCLUSION: This work demonstrates it is feasible to develop a spreadsheet-formula-based natural language processing tool to accurately extract information about drug overdose deaths from narratives; for most attributes, a rule-based search performs well or better than deep learning. The Excel Extractor has the potential to streamline data abstraction for epidemiologists gathering data about drug overdose deaths.

References

  1. 1.Omaki, E., Restrepo, F., Shields, W. C., & Abrahams, A. (2026). Natural language processing tool for extracting information about opioid overdoses in the USA from case narratives in the violent death reporting system. *Injury Prevention: Journal of the International Society for Child and Adolescent Injury Prevention*. Advance online publication. https://doi.org/10.1136/ip-2024-045314
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service