Extracting and Classifying Drug Discontinuations From Estonian Electronic Health Records: Development and Validation Study
Clinical Snapshot
PICO Framework
| P — Population | Adults with prescriptions for statins or antidiabetic medications drawn from a 10% population sample of Estonian electronic health records (EHR), 2012–2019 |
| I — Intervention | Application of large language models (LLMs: Llama 3.1-70B and GPT-4o) to free-text clinical anamneses for extraction and classification of drug discontinuation events |
| C — Comparator | Manual clinician review of 100 randomly selected cases per drug group (gold-standard validation set); no head-to-head comparison with traditional NLP methods in the primary analysis |
| O — Outcomes | Primary: precision of phrase extraction and reason extraction; weighted F1-scores for classification of discontinuation reasons and identification of who initiated discontinuation. Secondary: characterisation of discontinuation patterns (reason categories, initiator) for statins and antidiabetic drugs |
Bottom Line
This Estonian development and validation study demonstrates that large language models can extract and classify medication discontinuation reasons from free-text clinical notes with high precision (0.93–0.98) and acceptable F1-scores for reason classification (0.81–0.84). Adverse reactions dominated discontinuation events for both statins and antidiabetic drugs. The work is methodologically innovative and addresses a genuine gap in pharmacovigilance — the systematic capture of unstructured EHR data on why patients stop medications. However, several limitations temper enthusiasm: the validation sample is small (n=100 per drug group) with no confidence intervals reported; inter-rater reliability of the gold standard is undescribed; initiator classification performance is modest (F1: 0.64–0.78); and the pipeline is language-specific to Estonian, limiting direct transferability. For Australian clinicians and health informaticians, the conceptual approach is compelling and warrants adaptation to English-language EHR systems, but the tool itself cannot be adopted without revalidation. The study is best interpreted as a proof-of-concept for scalable, LLM-driven pharmacovigilance from routine clinical notes — a promising signal that merits larger, multi-site replication with rigorous reporting of uncertainty estimates.
Key Findings
P Value: Not reported
Effect Size: Phrase extraction precision: 0.93–0.98; reason extraction precision: 0.95–0.96; discontinuation reason classification weighted F1: 0.81–0.84; initiator classification weighted F1: 0.64–0.78
Primary Outcome: LLM-based extraction and classification of drug discontinuation events from Estonian free-text EHRs. Extraction yielded 625 antidiabetic and 233 statin discontinuation cases. Adverse reactions were the most common reason: 70% (163/233) of statin discontinuations and 44.8% (280/625) of antidiabetic discontinuations.
Nnt Or Sensitivity: Sensitivity (recall) not reported in the abstract. Weighted F1-scores serve as the primary composite performance metric. Initiator classification F1 of 0.64–0.78 indicates meaningful misclassification risk for this clinically important task.
Confidence Interval: Not reported for any metric
Clinical Application
The pipeline is technically feasible at scale within a population-based EHR system. However, deployment requires: (1) language-specific LLM fine-tuning or prompting; (2) data governance frameworks for processing identifiable clinical notes through LLMs (particularly proprietary models such as GPT-4o); (3) clinician involvement in taxonomy development and ongoing validation. The open-source Llama model offers a privacy-preserving alternative to cloud-based APIs. Direct applicability to Australian practice is limited. The pipeline is language-specific and would require complete revalidation for English-language Australian EHRs. However, the conceptual framework is highly relevant: Australian primary care EHRs (Best Practice, Medical Director) contain substantial free-text data on medication changes that are not captured in structured PBS dispensing records. The RACGP supports medication review and deprescribing frameworks (e.g., MedsCheck, HMR) where automated discontinuation reason capture could add value. The TGA's pharmacovigilance programme relies heavily on voluntary adverse event reporting; an LLM-based pipeline applied to Australian EHRs could complement spontaneous reporting. PBS data linkage studies (e.g., via AIHW) could provide the structured prescription backbone analogous to the Estonian prescription registry. Statins and antidiabetic agents are among the highest-volume PBS-subsidised drug classes, making this domain directly relevant. Data sovereignty and privacy considerations under the Australian Privacy Act 1988 and My Health Record Act 2012 would need careful navigation before deploying proprietary LLMs on identifiable clinical notes. Patients with chronic disease prescriptions (statins, antidiabetic agents) whose medication discontinuation events are documented in free-text clinical notes within EHR systems. Most directly applicable to healthcare systems using Estonian-language EHRs.
Abstract
BACKGROUND: Drug adherence is crucial for chronic disease management, yet treatment discontinuation remains common due to factors such as side effects, inefficacy, or cost. These reasons are often recorded only in free-text clinical notes, making large-scale analysis difficult. While large language models (LLMs) can interpret such unstructured data more effectively than traditional natural language processing methods, few studies have systematically categorized reasons for discontinuation or identified whether the decision was initiated by the patient or the clinician, especially in low-resource languages such as Estonian. OBJECTIVE: This study aimed to assess the ability of LLMs to extract and classify reasons for drug discontinuation and identify who initiated it using Estonian electronic health records and characterize the observed discontinuation patterns and initiators for statins and antidiabetic medications. METHODS: We combined prescription data with free-text anamneses from a 10% sample of the Estonian population (2012-2019). LLMs (Llama 3.1-70B and GPT-4o) were applied to extract discontinuation phrases and reasons, classify them into a clinician-developed taxonomy, and identify who discontinued the treatment. Performance was evaluated on 100 randomly chosen cases per drug group. RESULTS: Extraction yielded 625 antidiabetic drug and 233 statin discontinuation cases. Validation confirmed a precision of 0.93 to 0.98 for extracting phrases and 0.95 to 0.96 for extracting reasons. Classification of discontinuation reasons achieved weighted F1-scores of 0.81 to 0.84, whereas classification of who initiated discontinuation achieved weighted F1-scores of 0.64 to 0.78. Adverse reactions were the most frequent reason overall, accounting for 70% (163/233) of statin discontinuations and 44.8% (280/625) of antidiabetic drug discontinuations. Regarding antidiabetic drugs, treatment inefficacy and contraindications were more common. Patients more often stopped due to adverse reactions or nonmedical reasons, whereas physicians more often initiated discontinuation for contraindications. CONCLUSIONS: LLMs can accurately extract and classify medication discontinuation reasons and show variable performance in identifying discontinuation initiators in Estonian clinical narratives. Both local and proprietary models showed promising results, enabling scalable analyses that complement structured health records. This demonstrates the potential of LLMs to unlock information from clinical notes, turning this underused electronic health record component into a valuable resource for monitoring treatment patterns and detecting adverse event signals.
References
- 1.Šuvalov, H., Umov, N., Malk, M., Haug, M., Laur, S., Oja, M., Tamm, S., Reisberg, S., Vilo, J., & Kolde, R. (2026). Extracting and classifying drug discontinuations from Estonian electronic health records: Development and validation study. Journal of Medical Internet Research. https://doi.org/10.2196/86183
Related Research
CPT: pharmacometrics & systems pharmacology
Structured Schemas for Provenance-Rich, LLM-Assisted QSP Model Calibration
2 Aug 2026
Journal of pharmacokinetics and pharmacodynamics
Diffusion models for virtual populations and pharmacometric simulations
1 Aug 2026
Journal of chemical information and modeling
HyperDC: A Non-Uniform Hypergraph Framework for Dual- and Higher-Order Drug Combination Recommendation Across Diverse Complex Diseases
29 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service