Hepatology e-consult responses generated by artificial intelligence demonstrate accuracy but require human oversight.
Clinical Snapshot
PICO Framework
| P — Population | Adult patients referred for hepatology electronic consultations (e-consults) at the University of California San Francisco (UCSF) between January and March 2025, covering categories including abnormal liver function tests, hepatitis B, and abnormal imaging |
| I — Intervention | LiVersa-generated e-consult draft responses — a customised large language model (LLM) trained for liver disease clinical decision support |
| C — Comparator | Original hepatologist-authored e-consult responses, evaluated by three independent hepatologists and an LLM-as-a-judge (OpenAI-o1) using a 12-item rubric |
| O — Outcomes | Primary: clinical equivalence, immediate usability, comprehensiveness, presence of misleading or incorrect information, and risk of patient harm; Secondary: equivalence between human and LLM-based reviewers assessed via two one-sided tests (TOST) |
Bottom Line
This single-centre cross-sectional study from UCSF evaluated LiVersa, a customised large language model, as a drafting tool for hepatology electronic consultations across 61 cases. While 72% of drafts were rated as reasonable starting points and 83% provided appropriate case-specific recommendations, 10% contained misleading or incorrect information and 3.4% posed a risk of severe patient harm — rates that unequivocally preclude unsupervised clinical deployment. LiVersa performed comparably to hepatologist responses in length and verbosity but fell short on clinical equivalence, immediate usability, and comprehensiveness. A notable finding was the substantial divergence between human and LLM-based reviewer ratings, particularly for harm identification (20% vs. 67%), raising important questions about reviewer calibration and rubric standardisation. The study is methodologically limited by its small sample, single-centre design, absence of confirmed reviewer blinding, unreported inter-rater reliability, and lack of patient outcome data. For Australian clinicians and health services, this technology remains investigational. Any future implementation would require local validation, TGA SaMD regulatory compliance, alignment with GESA and RACGP guidelines, and mandatory human oversight workflows. The study's most clinically actionable message is that LLM-generated specialist consultation drafts require systematic expert review before reaching referring clinicians — a conclusion that should temper enthusiasm for rapid adoption.
Key Findings
P Value: Word count comparison p=0.47; verbosity p=0.44; TOST equivalence between human and LLM-as-a-judge reviewers p<0.05 for accuracy, precision, and comprehensiveness domains
Effect Size: LiVersa drafts rated as clinically equivalent by human reviewers: 48%; by LLM-as-a-judge: 27%. Potentially harmful ratings: 20% (human) vs. 67% (LLM-as-a-judge). No significant difference in word count (284 vs. 264 words, p=0.47) or verbosity (24 vs. 25 words/sentence, p=0.44)
Primary Outcome: 72% of LiVersa-generated e-consult drafts were rated as reasonable starting points by human hepatologist reviewers; 83% provided appropriate case-specific recommendations; 10% contained misleading or incorrect information; 3.4% posed a risk of severe patient harm
Nnt Or Sensitivity: Not applicable as a primary metric; as a quality/safety proxy: 1 in 10 LiVersa drafts contained misleading/incorrect information; 1 in 29 posed risk of severe harm — clinically relevant error rates requiring mandatory human review before clinical use
Confidence Interval: Not reported for primary performance metrics in the abstract; TOST mean differences for reviewer equivalence on accuracy, precision, and comprehensiveness: 0.026–0.029
Clinical Application
Technically feasible as a draft-generation tool to reduce hepatologist administrative burden, provided robust human oversight workflows are mandated. The 10% misleading information rate and 3.4% severe harm risk make unsupervised deployment clinically unacceptable. Implementation would require: validated rubric-based quality assurance, clear clinician accountability frameworks, version-controlled LLM deployment, and ongoing performance monitoring. Integration with existing electronic medical record e-consult platforms would require institutional IT infrastructure investment. No equivalent customised hepatology LLM tool is currently TGA-approved or in widespread clinical use in Australia. Australian hepatology e-consult infrastructure varies significantly across jurisdictions, with the My Health Record system and state-based telehealth platforms providing the primary digital consultation frameworks. GESA (Gastroenterological Society of Australia) and RACGP guidelines would need to be incorporated into any locally adapted LLM training corpus. The Australian hepatology disease burden differs from the UCSF cohort — notably higher rates of metabolic dysfunction-associated steatotic liver disease (MASLD), distinct Indigenous liver disease patterns, and different hepatitis B epidemiology in migrant communities — all of which would require local model customisation. PBS-listed therapies for hepatitis B (entecavir, tenofovir) and hepatitis C (pan-genotypic DAAs) differ from US formulary, and any LLM-generated recommendation referencing treatment would require Australian formulary alignment. TGA oversight of software as a medical device (SaMD) regulations would apply to any clinical deployment. This study does not provide sufficient evidence to support clinical implementation in Australian practice without local validation. Hepatology specialist services operating electronic consultation (e-consult) programmes, particularly in high-volume academic or tertiary referral centres where specialist access is constrained. Most directly applicable to hepatology e-consults involving abnormal liver function tests, hepatitis B management queries, and abnormal hepatic imaging interpretation.
Abstract
BACKGROUND: Electronic consultations (e-consults) improve specialist access but burden providers. We developed LiVersa, a customized large language model (LLM) for liver diseases. We evaluated its performance in drafting hepatology e-consult responses and the equivalence between human and machine reviewers. METHODS: LiVersa-generated responses for hepatology e-consults answered at the University of California San Francisco (UCSF) from January to March 2025. Using a 12-item rubric, 3 independent hepatologists and "LLM-as-a-judge" (OpenAI-o1) evaluated drafts against original responses. We tested equivalence between human reviewers and "LLM-as-a-judge" using two one-sided tests (TOST). RESULTS: Among 61 e-consults, the most common categories were abnormal liver function tests (34%), hepatitis B (23%), and abnormal imaging (21%). LiVersa drafts demonstrated no differences from hepatologist responses in word count (284 vs. 264, p=0.47) and verbosity (24 vs. 25 words per sentence, p=0.44). Human reviewers rated 72% of drafts as reasonable starting points and 83% as providing appropriate case-specific recommendations; 10% contained misleading/incorrect information, and 3.4% posed a risk of severe harm. LiVersa performed better at avoiding misleading information and extraneous suggestions but scored lower on clinical equivalence, immediate usability, and comprehensiveness. LLM-based reviewers were more stringent than human reviewers, rating fewer drafts as clinically equivalent (27% vs. 48%) and more as potentially harmful (67% vs. 20%), with agreement on accuracy, precision, and comprehensiveness (mean difference 0.026-0.029; TOST p<0.05). CONCLUSIONS: Customized LLMs like LiVersa show promise for e-consult drafting but require human oversight. LLM-as-a-judge was more conservative than humans, supporting its role in rapid quality assurance during model updates.
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service