Evaluating the diagnostic accuracy of ChatGPT-5 in the interpretation of ROTEM data
MINERVA ANESTESIOLOGICA, 2026 (SCI-Expanded, Scopus)
- Publication Type: Article / Article
- Publication Date: 2026
- Doi Number: 10.23736/s0375-9393.26.20024-x
- Journal Name: MINERVA ANESTESIOLOGICA
- Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus, EMBASE, MEDLINE
- Ondokuz Mayıs University Affiliated: Yes
Abstract
BACKGROUND: Rotational thromboelastometry (ROTEM) is widely used for goal-directed hemostatic management in cardiac and liver surgery. However, its interpretation requires algorithm-based expertise and remains subject to inter-observer variability. Large language models (LLMs) may assist clinicians in structured interpretation of viscoelastic data. METHODS: In this multicenter, retrospective validation study, 93 structured ROTEM-based clinical vignettes derived from 72 adult patients undergoing cardiac surgery or liver transplantation were evaluated. Two experienced anesthesiologists independently interpreted each vignette according to the G & ouml;rlinger algorithm, with adjudication in cases of disagreement. ChatGPT-5 was prompted using a standardized algorithm-based framework and generated binary diagnostic and treatment decisions for 11 predefined items. Agreement was assessed using Cohen's kappa (x), and diagnostic performance metrics were calculated. RESULTS: Inter-expert raw agreement was 86.5%. For the primary outcome (need for hemostatic treatment), agreement between ChatGPT-5 and expert consensus was high (x 0.703; 95% CI 0.526-0.853), with an observed agreement of 88.2%. Agreement was strongest for antifibrinolytic therapy (x 0.888) and fibrinogen replacement (x 0.881), and moderate for PCC/FFP transfusion (x 0.671). Diagnostic accuracy ranged from 0.699 (coagulopathy) to 0.957 (hyperfibrinolysis and fibrinogen deficiency). Sensitivity was highest for heparin effect (1.000) and lowest for protamine overdose (0.222). Performance was less consistent in platelet-related abnormalities. CONCLUSIONS: ChatGPT-5 demonstrated substantial agreement with expert ROTEM interpretation, particularly for conditions characterized by distinct viscoelastic patterns such as hyperfibrinolysis and fibrinogen deficiency. LLM-assisted interpretation may serve as a decision-support tool in ROTEM-guided bleeding management, although expert oversight remains essential in complex or multifactorial scenarios.