Vollständiger Abstract
Worum geht es in dieser Arbeit?
<h4>Background</h4>Clinical AI decision support is being introduced into nursing practice; however, existing large language models (LLMs) demonstrate only moderate accuracy on complex clinical tasks, raising questions about the level of accuracy required for safe clinical use across varying levels of clinician experience and task complexity.<h4>Objective</h4>The aim of the study is to develop an empirically calibrated simulation model of human-AI reliance and error in nursing decision-making and estimate the AI accuracy required to achieve specified error rate targets.<h4>Methods</h4>A linear reliance model with coefficients for AI accuracy (A), clinician experience (E), and task complexity (C) was calibrated using weighted least squares against 9 empirical data points from 3 independent randomized experiments on AI-assisted decision-making (N=3502). Predicted error was computed as reliance×(1-A) across a 27-cell factorial design. Study-level bootstrap (2000 iterations) quantified calibration uncertainty. To contextualize the simulation's operating range, the accuracy of contemporary general-purpose LLMs on complex clinical tasks was drawn from published benchmarks.<h4>Results</h4>Calibration placed βA at 0.201 (bootstrap 95% CI 0.023-0.234; P(βA>0)>.99). For the novice×high-complexity combination, the minimum AI accuracy values required to achieve error rates <10% and<20% were 0.89 and 0.78, respectively (bootstrap 95% CIs 0.88-0.90 and 0.75-0.79). At the moderate accuracy levels currently reported for general-purpose LLMs on complex clinical tasks (approximately 0.5-0.7), the model predicts error rates of roughly 26% to 41% in this high-risk condition.<h4>Conclusions</h4>In this model, keeping predicted error rates below a stringent target (<10%) for high-complexity nursing decision support by novice clinicians requires AI accuracy of at least approximately 0.89, a level that current general-purpose LLMs may not reliably reach on complex clinical tasks. Because the model is calibrated on nonnursing reliance data, these thresholds are illustrative model outputs, not nursing-derived empirical standards. The Athreshold framework provides a decision-theoretic tool for evaluating the minimum AI accuracy requirement by user-and-task profile. Behavioral validation in nursing contexts remains an essential next step. Because the framework is independent of any specific model, it remains applicable as AI systems improve.
Abstract: PubMed · Datensatz
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Nicht angegeben
- Quelle
- CrossRef Listing of Deleted DOIs
- Publikation
- 2000-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0849-6757
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
(2000). 10.1016/0967-0653(95)99590-n. CrossRef Listing of Deleted DOIs. https://doi.org/10.2196/99590
Kontext