A review of machine learning in toxicology: current practices and reporting gaps
Clinical Snapshot
PICO Framework
| P — Population | Published literature applying machine learning (ML) methods in toxicological risk assessment, indexed in PubMed or published in Computational Toxicology (2022–2024) |
| I — Intervention | Machine learning and artificial intelligence approaches applied to toxicological data (including in vitro, in vivo, and computational toxicology contexts) |
| C — Comparator | No formal comparator; the review benchmarks current ML reporting practices against established methodological and reporting standards |
| O — Outcomes | Prevalence and quality of reporting practices for ML methods in toxicology, including internal validation, model interpretability, handling of missing data, and availability of data and code |
Bottom Line
This descriptive literature review from TU Dortmund provides a timely audit of how machine learning methods are reported in toxicology publications from 2022 to 2024. The headline finding — that roughly half of ML papers in toxicology fail to use any model interpretability method, and that reporting of missing data handling and code/data availability is frequently absent — is concerning for reproducibility and regulatory confidence. The review is methodologically modest: it draws on only two literature sources, applies no formal quality assessment tool, and provides no uncertainty estimates around its descriptive proportions. It does not evaluate whether the ML models actually work, only how they are reported. Nevertheless, its value lies in establishing a baseline for the field and identifying specific, addressable reporting gaps. For senior clinicians and regulatory scientists, the practical implication is clear: ML-derived toxicity predictions entering drug development pipelines, chemical risk assessments, or regulatory submissions should be scrutinised for interpretability, validation rigour, and data transparency. Journals, funders, and regulatory bodies — including the TGA and AICIS in Australia — should consider mandating structured ML reporting checklists analogous to TRIPOD or CONSORT. This review is best read as a call to action for the field rather than definitive evidence.
Key Findings
P Value: Not applicable; descriptive review without hypothesis testing
Effect Size: Approximately 50% of reviewed papers employed model interpretation methods to address 'black box' predictions; frequent gaps identified in reporting of missing value handling and availability of data and code (specific proportions not extractable from abstract alone)
Primary Outcome: Characterisation of ML reporting practices in toxicology literature (2022–2024): prevalence of internal validation, model interpretability methods, missing data handling, and data/code availability
Nnt Or Sensitivity: Not applicable; this is a methodological review, not a clinical intervention or diagnostic accuracy study. No NNT, sensitivity, specificity, or hazard ratio is calculable.
Confidence Interval: Not reported; no confidence intervals provided for descriptive frequency estimates
Clinical Application
The review's findings are directly actionable for researchers and journal editors seeking to improve ML reporting standards. Implementation of recommended reporting practices (e.g., mandatory data/code availability, explicit missing data handling, use of interpretability methods) is feasible at low cost. Adoption would require editorial policy changes at toxicology journals and updated author guidelines. In Australia, the Australian Industrial Chemicals Introduction Scheme (AICIS, formerly NICNAS) and the Therapeutic Goods Administration (TGA) increasingly encounter ML-based toxicological assessments in chemical notification and preclinical drug dossiers. The reporting gaps identified — particularly around model interpretability and data availability — are directly relevant to regulatory submission quality. The RACGP and clinical pharmacology community would benefit from improved transparency in ML-derived toxicity predictions used to inform prescribing safety data. Australia's participation in OECD chemical safety frameworks means that any international movement toward standardised ML reporting (analogous to OECD QSAR guidance documents) would have direct regulatory implications domestically. The TGA's adoption of ICH guidelines for preclinical safety (ICH S7A/S7B) may need to be updated to address ML-specific validation and reporting requirements. Toxicologists, computational chemists, regulatory scientists, and clinical pharmacologists who develop, evaluate, or apply ML-based predictive models for toxicological risk assessment. Relevant to preclinical drug development teams, chemical safety assessors, and academic researchers publishing in the field.
Abstract
In recent years, machine learning and artificial intelligence approaches have been increasingly applied in the context of toxicological risk assessment. Many published overview, review, and comment papers discuss advantages, disadvantages, success stories, and open challenges for the application of machine learning models in toxicology. Machine learning methods using information from in vitro experiments can help to avoid animal experiments, thus allowing for larger numbers of experiments to be conducted. Drawbacks of machine learning models are the lack of mechanistic interpretability and the need for large amounts of high-quality data. In this work, we present a literature review of papers indexed in PubMed or published in the journal Computational Toxicology in the years 2022 to 2024, to assess the usage of machine learning methods in toxicology as well as the practices in reporting of methods and corresponding results. We do not address the suitability or the performance of methods, which is impossible to assess objectively without reanalysis on raw data, but focus on common practices and gaps in reporting. Major results are that many different machine learning methods are used in toxicology, often with appropriate internal validation. However, in only half of the cases, interpretation methods are used to address the problem that these models often make predictions as a black box. Moreover, there are very frequent gaps in reporting, in particular related to handling of missing values, and availability of data and code. Thus, this review can serve as a starting point for further tailored methodological research and guidance.
References
- 1.Kappenberg, F., Stolte, M., Sauer, L., Duda, J. C., Lau, M., Schürmeyer, L., Zhou, H., Schwender, H., Schorning, K., & Rahnenführer, J. (2026). A review of machine learning in toxicology: current practices and reporting gaps. Archives of Toxicology. https://doi.org/10.1007/s00204-025-03987-2
Related Research
Cardiovascular toxicology
Mechanisms of Cardiovascular Toxicity Induced by Silver Nanoparticles: A Systematic Review of Preclinical Evidence
13 July 2026
JMIR infodemiology
Self-Reported Tianeptine Experiences on Reddit: Natural Language Processing-Assisted Qualitative Study
8 July 2026
Journal of chemical information and modeling
BioTD: An Online Database of Biotoxins
23 June 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service