Structured Schemas for Provenance-Rich, LLM-Assisted QSP Model Calibration
Clinical Snapshot
PICO Framework
| P — Population | Quantitative systems pharmacology (QSP) modellers working with pancreatic ductal adenocarcinoma (PDAC) models requiring literature-based parameter calibration |
| I — Intervention | MAPLE (Model-Aware Parameterization from Literature Evidence) — a structured schema-based framework using validated provenance-tracking workflows to assist QSP model calibration from literature |
| C — Comparator | Implicit comparison to conventional manual curation and unstructured large language model (LLM)-assisted extraction (no formal concurrent control group) |
| O — Outcomes | Accuracy and provenance completeness of extracted calibration targets; frequency of automated error detection and retry; proportion of parameters supported by multiple independent sources; extent of modeller revision required across SubmodelTarget and CalibrationTarget schemas |
Bottom Line
MAPLE is a thoughtfully designed methodological framework that addresses a genuine and underappreciated problem in quantitative systems pharmacology: the inconsistent documentation of calibration data provenance and the risk of hallucinated values when large language models are used for literature extraction. The two-schema architecture — separating isolated submodel experiments from full clinical calibration targets — is conceptually sound, and the automated validators (DOI resolution, source snippet matching, executable code checking) represent a meaningful quality-control advance over unstructured workflows. However, this paper is a single-group, single-disease-area proof-of-concept with no comparator arm, no inferential statistics, and no downstream validation of model predictive performance. The 65% modeller revision rate for forward-model choices and 100% revision rate for source relevance assessments confirm that MAPLE augments rather than replaces expert judgement. Clinicians and pharmacometricians should regard this as a promising early-stage framework warranting independent multi-group validation across diverse disease areas before adoption as a standard practice. The open schema design facilitates such validation. For Australian regulatory and HTA contexts, the provenance architecture is well-aligned with TGA and PBAC transparency expectations, but formal regulatory endorsement awaits broader evidence.
Key Findings
P Value: Not reported — no inferential statistics presented
Effect Size: Not applicable — no comparative effect size reported; descriptive metrics only
Primary Outcome: Successful extraction and curation of 37 SubmodelTargets and 45 CalibrationTargets for a PDAC QSP model using the MAPLE framework, with full provenance (direct source quotes and verified citations) for every extracted value
Nnt Or Sensitivity: 50 automated retries triggered before human review; modeller revised forward-model choices in 65% of SubmodelTargets, priors in 46%, and source relevance assessments in 100% of files; 11/19 parameters (58%) supported by more than one independent source
Confidence Interval: Not reported
Clinical Application
Conceptually feasible for well-resourced academic and industry modelling groups with access to LLM APIs and schema validation tooling. Implementation requires modeller investment in schema familiarisation and context preparation. Scalability to large model libraries or resource-limited settings is undemonstrated. The 100% modeller revision rate for source relevance indicates that full automation remains unrealistic and sustained expert oversight is essential. Directly relevant to Australian pharmaceutical industry and academic QSP modelling groups, including those supporting TGA regulatory submissions and PBAC health technology assessments where model transparency and reproducibility are increasingly scrutinised. The Therapeutic Goods Administration (TGA) and the Pharmaceutical Benefits Advisory Committee (PBAC) place growing emphasis on model provenance and reproducibility in submissions involving mechanistic pharmacokinetic-pharmacodynamic and systems models. MAPLE's provenance-recording architecture aligns with these regulatory expectations. Australian academic centres with QSP programmes (e.g., University of Melbourne, Monash University, UNSW) and contract research organisations supporting oncology drug development would be the primary adopters. No PBS or RACGP guideline implications arise directly from this methodological paper. Pharmacometricians, systems pharmacologists, and quantitative modellers engaged in QSP model development for drug development, regulatory submission, or translational research — particularly those working in oncology or other complex disease areas requiring multi-scale parameter calibration from heterogeneous literature
Abstract
Quantitative systems pharmacology (QSP) models require calibration data from literature, yet manual curation is inconsistently documented and large language model (LLM) extraction can hallucinate values and fabricate citations. We present MAPLE (Model-Aware Parameterization from Literature Evidence), which uses structured validation schemas as a collaboration interface between LLMs and modelers. Two schemas span two scales: the SubmodelTarget schema for isolated experiments constraining individual parameters, and the CalibrationTarget schema for clinical and in vivo endpoints constraining the full model. Both separate data extraction from modeling decisions, recording every value with full provenance. Targeted validators catch characteristic LLM errors by matching values to source snippets, resolving DOIs, and executing code. For a pancreatic ductal adenocarcinoma QSP model, we used MAPLE to extract and curate 37 SubmodelTargets and 45 CalibrationTargets. Before any human review, the validators triggered 50 automated retries; every value carries a direct quote from its source and a verified citation; and 11 of 19 parameters are supported by more than one independent source. The LLM drafted usable forward models and code from context, while the modeler supplied the context and scientific judgment it cannot infer, revising forward-model choices in 65% of SubmodelTargets, priors in 46%, and source relevance in all files. This evaluation covers one model in one disease area, by a single group, so it characterizes the framework rather than establishing how broadly it generalizes. MAPLE records the modeler's reasoning in a form that can be re-run and independently checked, so it is not lost when the modeling team changes.
References
- 1.Eliason, J., & Popel, A. S. (2026). Structured schemas for provenance-rich, LLM-assisted QSP model calibration. CPT: Pharmacometrics & Systems Pharmacology. https://doi.org/10.1093/bioadv/vbae194
Related Research
Journal of pharmacokinetics and pharmacodynamics
Diffusion models for virtual populations and pharmacometric simulations
1 Aug 2026
Journal of chemical information and modeling
HyperDC: A Non-Uniform Hypergraph Framework for Dual- and Higher-Order Drug Combination Recommendation Across Diverse Complex Diseases
29 July 2026
Journal of chemical information and modeling
A Unified Molecular Graph and Protein Language Model Framework for Predicting Human Drug-Hormone Receptor Interactions with Structure-Aware Validation
28 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service