<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.0 20040830//EN" "journalpublishing.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="2.0" xml:lang="en" article-type="research-article"><front><journal-meta><journal-id journal-id-type="nlm-ta">JMIR Nursing</journal-id><journal-id journal-id-type="publisher-id">nursing</journal-id><journal-id journal-id-type="index">33</journal-id><journal-title>JMIR Nursing</journal-title><abbrev-journal-title>JMIR Nursing</abbrev-journal-title><issn pub-type="epub">2562-7600</issn><publisher><publisher-name>JMIR Publications</publisher-name><publisher-loc>Toronto, Canada</publisher-loc></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">v9i1e99590</article-id><article-id pub-id-type="doi">10.2196/99590</article-id><article-categories><subj-group subj-group-type="heading"><subject>Original Paper</subject></subj-group></article-categories><title-group><article-title>Theoretical Exploration of Error Thresholds for Clinical AI Decision Support in Nursing: Exploratory Simulation Study Grounded in Human-AI Reliance Data</article-title></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name name-style="western"><surname>Tajima</surname><given-names>Hiroyuki</given-names></name><degrees>PhD</degrees><xref ref-type="aff" rid="aff1"/></contrib></contrib-group><aff id="aff1"><institution>Faculty of Nursing, Shumei University</institution><addr-line>1-1 Daigaku-cho</addr-line><addr-line>Yachiyo</addr-line><addr-line>Chiba</addr-line><country>Japan</country></aff><contrib-group><contrib contrib-type="editor"><name name-style="western"><surname>Borycki</surname><given-names>Elizabeth</given-names></name></contrib><contrib contrib-type="editor"><name name-style="western"><surname>Sarvestan</surname><given-names>Javad</given-names></name></contrib></contrib-group><contrib-group><contrib contrib-type="reviewer"><name name-style="western"><surname>Quindroit</surname><given-names>Paul</given-names></name></contrib><contrib contrib-type="reviewer"><name name-style="western"><surname>Kim</surname><given-names>Aeran</given-names></name></contrib></contrib-group><author-notes><corresp>Correspondence to HiroyukiTajima, PhD, Faculty of Nursing, Shumei University, 1-1 Daigaku-cho, Yachiyo, 276-0003, Chiba, Japan, 81 47-488-2111; <email>tajima@mailg.shumei-u.ac.jp</email></corresp></author-notes><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>5</day><month>8</month><year>2026</year></pub-date><volume>9</volume><elocation-id>e99590</elocation-id><history><date date-type="received"><day>27</day><month>04</month><year>2026</year></date><date date-type="rev-recd"><day>16</day><month>07</month><year>2026</year></date><date date-type="accepted"><day>17</day><month>07</month><year>2026</year></date></history><copyright-statement>&#x00A9; Hiroyuki Tajima. Originally published in JMIR Nursing (<ext-link ext-link-type="uri" xlink:href="https://nursing.jmir.org">https://nursing.jmir.org</ext-link>), 5.8.2026. </copyright-statement><copyright-year>2026</copyright-year><license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (<ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link>), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR Nursing, is properly cited. The complete bibliographic information, a link to the original publication on <ext-link ext-link-type="uri" xlink:href="https://nursing.jmir.org/">https://nursing.jmir.org/</ext-link>, as well as this copyright and license information must be included.</p></license><self-uri xlink:type="simple" xlink:href="https://nursing.jmir.org/2026/1/e99590"/><abstract><sec><title>Background</title><p>Clinical AI decision support is being introduced into nursing practice; however, existing large language models (LLMs) demonstrate only moderate accuracy on complex clinical tasks, raising questions about the level of accuracy required for safe clinical use across varying levels of clinician experience and task complexity.</p></sec><sec><title>Objective</title><p>The aim of the study is to develop an empirically calibrated simulation model of human-AI reliance and error in nursing decision-making and estimate the AI accuracy required to achieve specified error rate targets.</p></sec><sec sec-type="methods"><title>Methods</title><p>A linear reliance model with coefficients for AI accuracy (A), clinician experience (E), and task complexity (C) was calibrated using weighted least squares against 9 empirical data points from 3 independent randomized experiments on AI-assisted decision-making (N=3502). Predicted error was computed as reliance&#x00D7;(1&#x2212;<italic>A</italic>) across a 27-cell factorial design. Study-level bootstrap (2000 iterations) quantified calibration uncertainty. To contextualize the simulation&#x2019;s operating range, the accuracy of contemporary general-purpose LLMs on complex clinical tasks was drawn from published benchmarks.</p></sec><sec sec-type="results"><title>Results</title><p>Calibration placed &#x03B2;<sub>A</sub> at 0.201 (bootstrap 95% CI 0.023&#x2010;0.234; <italic>P</italic>(&#x03B2;<sub>A</sub>&#x003E;0)&#x003E;.99). For the novice&#x00D7;high-complexity combination, the minimum AI accuracy values required to achieve error rates &#x003C;10% and&#x003C;20% were 0.89 and 0.78, respectively (bootstrap 95% CIs 0.88&#x2010;0.90 and 0.75&#x2010;0.79). At the moderate accuracy levels currently reported for general-purpose LLMs on complex clinical tasks (approximately 0.5&#x2010;0.7), the model predicts error rates of roughly 26% to 41% in this high-risk condition.</p></sec><sec sec-type="conclusions"><title>Conclusions</title><p>In this model, keeping predicted error rates below a stringent target (&#x003C;10%) for high-complexity nursing decision support by novice clinicians requires AI accuracy of at least approximately 0.89, a level that current general-purpose LLMs may not reliably reach on complex clinical tasks. Because the model is calibrated on nonnursing reliance data, these thresholds are illustrative model outputs, not nursing-derived empirical standards. The <italic>A</italic><sub>threshold</sub> framework provides a decision-theoretic tool for evaluating the minimum AI accuracy requirement by user-and-task profile. Behavioral validation in nursing contexts remains an essential next step. Because the framework is independent of any specific model, it remains applicable as AI systems improve.</p></sec></abstract><kwd-group><kwd>clinical decision support</kwd><kwd>automation bias</kwd><kwd>large language models</kwd><kwd>nursing informatics</kwd><kwd>Monte Carlo simulation</kwd><kwd>error threshold</kwd></kwd-group></article-meta></front><body><sec id="s1" sec-type="intro"><title>Introduction</title><p>The integration of AI, particularly large language models (LLMs), into nursing education and clinical practice has increased substantially in recent years [<xref ref-type="bibr" rid="ref1">1</xref>-<xref ref-type="bibr" rid="ref4">4</xref>]. Although these systems present considerable opportunities to augment clinical reasoning and decision-making, they simultaneously introduce novel risks associated with human reliance on AI-generated outputs [<xref ref-type="bibr" rid="ref5">5</xref>]. A critical concern is how accurate current AI systems must be for safe use in clinical tasks, particularly because clinical decision support systems can improve care but also introduce safety risks when their outputs are relied upon uncritically [<xref ref-type="bibr" rid="ref6">6</xref>-<xref ref-type="bibr" rid="ref8">8</xref>].</p><p>Research in human factors and automation science has long demonstrated that users often develop inappropriate levels of trust in automated systems, resulting in overreliance and consequent errors [<xref ref-type="bibr" rid="ref9">9</xref>,<xref ref-type="bibr" rid="ref10">10</xref>]. Parasuraman and Riley [<xref ref-type="bibr" rid="ref9">9</xref>] identified &#x201C;misuse&#x201D;&#x2014;overreliance on automation that produces failures in monitoring or detecting system errors, as a principal failure mode of human-automation interaction. Lee and See [<xref ref-type="bibr" rid="ref10">10</xref>] subsequently articulated a framework of appropriate reliance in which trust should be calibrated to a system&#x2019;s demonstrated capabilities. In health care specifically, systematic reviews have documented automation bias as a recurring safety concern in computerized clinical decision support, identifying task complexity, cognitive load, and verification demands as principal mediators of overreliance [<xref ref-type="bibr" rid="ref11">11</xref>-<xref ref-type="bibr" rid="ref13">13</xref>]. Despite the established relevance of these frameworks, their quantitative application to health care AI, and to nursing in particular, remains largely underexplored.</p><p>Recent empirical studies have begun to characterize human-AI reliance relationships in AI-assisted decision-making. Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>] conducted 3 preregistered randomized experiments (total n=3793) and reported that participants relied more on AI systems, as the systems&#x2019; stated and observed accuracy increased, with agreement fractions rising from approximately 0.75 at 60% accuracy to 0.82 at 95% accuracy in their first experiment. Lu and Yin [<xref ref-type="bibr" rid="ref15">15</xref>] reported a similar monotonic trend in a subsequent study (n=466, 13,980 trials). In clinical settings, K&#x00FC;cking et al [<xref ref-type="bibr" rid="ref16">16</xref>] demonstrated that AI recommendation correctness exerted a strong bidirectional influence on the diagnostic accuracy of 223 physicians and nurses assessing wound maceration; correct AI recommendations increased the odds of correct decisions approximately 10-fold (odds ratio [OR] 10.0; <italic>P</italic>&#x003C;.001), whereas incorrect recommendations significantly reduced accuracy. These findings indicate that automation complacency concerns extend to contemporary AI-based clinical decision support systems and motivate a theoretically grounded, quantitative investigation of the conditions under which AI-assisted clinical errors are most likely to occur.</p><p>However, a critical gap remains between the available empirical literature and clinical use decisions. Current studies characterize human-AI reliance in qualitative terms or across a limited range of accuracy levels; yet, they do not provide clinicians or administrators with a decision-theoretic tool for determining whether an AI system, given its measured accuracy, is appropriate for a specific clinical task and user group. Such a tool is increasingly needed because, although LLMs achieve high scores on standardized medical knowledge benchmarks [<xref ref-type="bibr" rid="ref17">17</xref>], their accuracy on more complex, open-ended clinical decision-making tasks remains only moderate [<xref ref-type="bibr" rid="ref18">18</xref>,<xref ref-type="bibr" rid="ref19">19</xref>], placing current systems within an operating range where the clinical safety implications are neither self-evident nor reliably inferable from the raw accuracy value alone.</p><p>To address this gap, this study developed a linear simulation model of human-AI reliance in nursing decision-making, empirically calibrated against 9 data points derived from 3 independent randomized experiments on AI-assisted decision-making (total N=3502). Error rates were computed as the product of predicted reliance and AI failure probability and evaluated across a 27-cell factorial design spanning 3 AI accuracy levels, 3 clinician experience levels, and 3 task complexity levels. This work has two primary contributions: (1) empirical calibration of the reliance model from 3 independent large-scale datasets, which avoids the circularity inherent in purely theoretical parameter selection; and (2) introduction of <italic>A</italic><sub>threshold</sub>, a decision-theoretic parameter that quantifies the minimum AI accuracy required to achieve a specified clinical error rate target for a specific user-and-task profile.</p><p>Clinical AI decision support is being introduced into nursing practice without quantitative guidance on the minimum AI accuracy required for safe use. The <italic>A</italic><sub>threshold</sub> framework presented here is expected to benefit nursing educators, clinical informaticists, hospital administrators evaluating AI deployment, and regulators developing accuracy standards for clinical AI decision support systems. This work is submitted to the JMIR Nursing theme issue on Artificial Intelligence (AI) in Nursing, addressing the in-scope topics of LLM use in settings where nurses provide care, comparisons of AI algorithm effectiveness in supporting decision-making among nurses, and strengths and limitations of AI technologies in nursing practice.</p></sec><sec id="s2" sec-type="methods"><title>Methods</title><sec id="s2-1"><title>Study Design</title><p>This study used a theoretical Monte Carlo simulation calibrated against published human-AI reliance data. The published human data were used to calibrate the model&#x2019;s accuracy-related parameters, while all other parameters not determined by calibration were specified a priori from theoretical considerations and tested through sensitivity analysis. To ensure that the simulation&#x2019;s operating range was ecologically plausible, the accuracy of contemporary general-purpose LLMs on complex clinical tasks was characterized using published benchmarks [<xref ref-type="bibr" rid="ref18">18</xref>,<xref ref-type="bibr" rid="ref19">19</xref>] rather than a separate in-house evaluation.</p></sec><sec id="s2-2"><title>Ethical Considerations</title><p>This study did not involve human participants or the use of patient data; it comprised a computational simulation calibrated against previously published, publicly available aggregate data. Accordingly, institutional review board approval was not required.</p></sec><sec id="s2-3"><title>Theoretical Framework and Model Formulation</title><sec id="s2-3-1"><title>Reliance Model</title><p>Human reliance on AI was modeled as a linear function of AI accuracy, clinician experience, and task complexity:</p><p>Reliance=&#x03B2;<sub>&#x2080;</sub>+&#x03B2;<sub>A</sub>&#x00B7;<italic>A</italic>+&#x03B2;<sub>E</sub>&#x00B7;<italic>E</italic>+&#x03B2;<sub>C</sub>&#x00B7;<italic>C</italic>+&#x03B2;<sub>AE</sub>&#x00B7;(<italic>A</italic>&#x00D7;<italic>E</italic>)+&#x03B2;<sub>AC</sub>&#x00B7;(<italic>A</italic>&#x00D7;<italic>C</italic>)+&#x03B5;&#x2081;(1)</p><p>where <italic>A</italic> denotes AI accuracy on [0, 1]; <italic>E</italic> signifies clinical experience normalized by dividing the years of experience by 10 (so that 0, 2, and 10 years map to 0.0, 0.2, and 1.0, respectively); <italic>C</italic> represents task complexity coded as 0.0 (low), 0.5 (moderate), and 1.0 (high); and &#x03B5;&#x2081;~<italic>N</italic>(0, &#x03C3;=0.05). To preserve its probabilistic interpretation, reliance was clipped to the interval [0, 1].</p><p>The linear form was chosen over more elaborate specifications (eg, a logistic trust mediator with a quadratic accuracy penalty) for 3 reasons. First, it is parsimonious, as it contains the minimum parameters needed to capture additive effects of accuracy, experience, and complexity with their 2-way interactions. Second, it can be directly calibrated against published reliance data using standard regression methods, whereas more elaborate specifications add extra parameters that cannot be reliably estimated from the available data. Third, empirical studies of human-AI reliance consistently reported approximately linear relationships between AI accuracy and reliance across the 0.5&#x2010;1.0 range [<xref ref-type="bibr" rid="ref14">14</xref>,<xref ref-type="bibr" rid="ref15">15</xref>], which empirically justifies the linear specification.</p></sec><sec id="s2-3-2"><title>Error Model</title><p>Error was modeled as the product of behavioral reliance and the AI failure rate with additive Gaussian noise:</p><p>Error=reliance&#x00D7;(1&#x2212;<italic>A</italic>)+&#x03B5;&#x2082;(2)</p><p>where &#x03B5;&#x2082;~<italic>N</italic>(0, &#x03C3;=0.02), with error subsequently clipped to [0, 1]. This specification assumes that an erroneous AI recommendation leads to a clinical error only to the extent that the clinician relies on it. <xref ref-type="fig" rid="figure1">Figure 1</xref> presents the conceptual model of human-AI interaction in nursing decision-making.</p><fig position="float" id="figure1"><label>Figure 1.</label><caption><p>Conceptual model of AI-assisted clinical decision-making. Inputs (AI accuracy A, clinician experience E, and task complexity C) feed the reliance equation with additive main effects and 2-way interactions (<italic>A</italic>&#x00D7;<italic>E</italic> and <italic>A</italic>&#x00D7;<italic>C</italic>). The accuracy-related coefficients &#x03B2;<sub>&#x2080;</sub> and &#x03B2;<sub>A</sub> were empirically calibrated against 9 data points drawn from 3 independent randomized experiments on AI-assisted decision-making (total N=3502). The remaining coefficients were fixed from prior theoretical considerations. Error is computed as reliance&#x00D7;(1&#x2212;<italic>A</italic>), so AI accuracy enters the error term both indirectly through reliance and directly through the (1&#x2212;<italic>A</italic>) factor. The predicted error is then compared against a specified target error rate. &#x03B2;<sub>0</sub> and &#x03B2;<sub>A</sub> were calibrated from 9 empirical reliance points [<xref ref-type="bibr" rid="ref14">14</xref>,<xref ref-type="bibr" rid="ref15">15</xref>]; &#x03B2;<sub>E</sub>, &#x03B2;<sub>C</sub>, &#x03B2;<sub>AE</sub>, and &#x03B2;<sub>AC</sub> were fixed from the literature.</p></caption><graphic alt-version="no" mimetype="image" position="float" xlink:type="simple" xlink:href="nursing_v9i1e99590_fig01.png"/></fig></sec></sec><sec id="s2-4"><title>Model Calibration</title><p>The <italic>A</italic>-related coefficients of the reliance model (&#x03B2;<sub>&#x2080;</sub> and &#x03B2;<sub>A</sub>) were calibrated using weighted least squares against 9 empirical data points derived from 3 independent randomized experiments on AI-assisted decision-making. Studies met 3 prespecified criteria: experimental manipulation of AI accuracy at &#x2265;2 levels in the 0.5&#x2010;1.0 range, reported behavioral measure of reliance (agreement fraction, switch fraction, or equivalent), and total study n&#x2265;100.</p><p>In total, 3 studies met all 3 criteria. In all 3 experiments, laypeople recruited via Amazon Mechanical Turk used a machine-learning classifier to predict the outcomes of speed-dating events (whether a given participant would want to meet their date again), and reliance was operationalized as the agreement (or switch) fraction between the participant&#x2019;s final decision and the model&#x2019;s prediction. Lu and Yin [<xref ref-type="bibr" rid="ref15">15</xref>] (experiment 2; n=466 crowdworkers; 13,980 trials) manipulated the designed AI accuracy (50% and 80%) and measured the final agreement rate. Individual-level data were obtained from authors&#x2019; publicly available dataset [<xref ref-type="bibr" rid="ref20">20</xref>] and reanalyzed to obtain worker-level reliance means at each accuracy level (<italic>A</italic>=0.50: n=252; mean 0.642, SD 0.139; <italic>A</italic>=0.80: n=214; mean 0.664, SD 0.134). Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>] (experiment 1, phase 1; n=1994; stated accuracy &#x2208; {60%, 70%, 90%, 95%}; agreement fraction with observed accuracy of 80% and experiment 3, phase 2; n=1042; observed accuracy &#x2208; {55%, 80%, 100%}; agreement fraction averaged across 2 stated accuracy levels) contributed 4 and 3 data points, respectively. Values were extracted from the published figures (Figures 3a and 5a) of Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>] by visual inspection. The 2 Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>] experiments used here (experiment 1, phase 1 and experiment 3, phase 2) together comprised 3036 of the 3793 participants reported across that study&#x2019;s 3 experiments; adding the 466 crowdworkers of Lu and Yin [<xref ref-type="bibr" rid="ref15">15</xref>] yields the total calibration sample of 3502. Experiments and phases that did not manipulate AI accuracy across at least 2 levels in the 0.5&#x2010;1.0 range were not used for calibration.</p><p>All calibration targets correspond to the clinician-experience value <italic>E</italic>=0 (laypeople crowdworkers) and the task complexity value <italic>C</italic>=0.5 (moderate complexity; speed-dating outcome prediction requires the integration of multifeature personal profiles). Calibration was performed by minimizing the sum of the squared residuals weighted by the square root of each study&#x2019;s sample size using the Nelder-Mead simplex algorithm in SciPy (version 1.13; NumFOCUS). The other coefficients (&#x03B2;<sub>E</sub>, &#x03B2;<sub>C</sub>, &#x03B2;<sub>AE</sub>, and &#x03B2;<sub>AC</sub>) were set from prior theoretical considerations and are described in the Parameter Specification section.</p><p>Calibration yielded &#x03B2;<sub>&#x2080;</sub>=0.471 (baseline reliance at <italic>A</italic>=0, <italic>E</italic>=0, <italic>C</italic>=0.5 with literature-fixed complexity contribution) and &#x03B2;<sub>A</sub>=0.201 (reliance increase per unit of AI accuracy). The calibrated model achieved a root-mean-square error (RMSE) of 0.054 across the 9 empirical points.</p><p>To evaluate the precision of the calibrated coefficients given the limited number of independent data sources, study-level bootstrap resampling was run with 2000 iterations. For each iteration, the 3 studies were sampled with replacement, and &#x03B2;<sub>&#x2080;</sub> and &#x03B2;<sub>A</sub> were recalibrated on the resulting target set. Leave-one-study-out cross-validation was also conducted by excluding each study in turn and recalibrating the remaining 2. The bootstrap distribution and leave-one-out results are presented in the Results section.</p></sec><sec id="s2-5"><title>Parameter Specification</title><p>The remaining 4 coefficients were specified a priori based on theoretical considerations and were not estimated from the calibration data. The coefficient for clinical experience in the reliance equation (&#x03B2;<sub>E</sub>=&#x2212;0.20) reflects two complementary mechanisms: (1) expertise development cultivates reliance on internalized clinical knowledge and pattern recognition rather than on external decision aids [<xref ref-type="bibr" rid="ref21">21</xref>,<xref ref-type="bibr" rid="ref22">22</xref>] and (2) empirical studies in clinical AI decision support have shown that longer work experience is directly associated with reduced susceptibility to incorrect AI recommendations (OR 1.89 for work experience in a recent study of 223 health care professionals [<xref ref-type="bibr" rid="ref16">16</xref>]). The coefficient for task complexity (&#x03B2;<sub>C</sub>=+0.20) operationalizes evidence that increased cognitive load is associated with reliance on decision aids [<xref ref-type="bibr" rid="ref9">9</xref>]. The 2 interaction coefficients (&#x03B2;<sub>AE</sub>=+0.10 and &#x03B2;<sub>AC</sub>=+0.10) represent, respectively, expertise-based attenuation and complexity-based amplification of accuracy effects. These values are theoretically motivated rather than empirically estimated. The specific magnitudes (|&#x03B2;|=0.10&#x2010;0.20) were chosen as conservative nominal values, deliberately set at or below the magnitude of the calibrated accuracy slope (|&#x03B2;<sub>A</sub>|&#x2248;0.20) so that these unvalidated assumptions would not dominate the empirically calibrated reliance-accuracy relationship; they are not presented as precise estimates. Because 9 aggregate calibration points cannot identify these 4 coefficients empirically, alternative specifications were not selected on the basis of statistical fit; instead, the influence of these values was bounded through the sensitivity analyses described below, which span a plausible range of alternative magnitudes.</p></sec><sec id="s2-6"><title>Simulation Conditions</title><p>In a Monte Carlo simulation, an outcome of interest is estimated by repeatedly drawing random values from specified probability distributions and averaging the results across many simulated trials; here, this approach propagated the small random variation in the modeled reliance and error terms into stable estimates of the mean predicted error rate for each condition. A full factorial design was used: 3 levels of AI accuracy (0.5, 0.7, and 0.9)&#x00D7;3 levels of clinical experience (0, 2, and 10 years)&#x00D7;3 levels of task complexity (low, moderate, and high)=27 conditions. Each condition was simulated with 10,000 Monte Carlo trials (total n=270,000 simulated decision trials). The pseudorandom number generator was initialized with seed 42 using the numpy default_rng interface for full reproducibility.</p></sec><sec id="s2-7"><title><italic>A</italic><sub>threshold</sub> Analysis</title><p>The minimum AI accuracy required to achieve a specified clinical error rate target (denoted as <italic>A</italic><sub>threshold</sub>) was computed analytically for each combination of clinical experience and task complexity. Given the linear reliance specification, the predicted mean error is a quadratic function of <italic>A</italic>:</p><p>Error(<italic>A</italic>|<italic>E</italic>, <italic>C</italic>)=[(&#x03B2;<sub>&#x2080;</sub>+&#x03B2;<sub>E</sub>&#x00B7;<italic>E</italic>+&#x03B2;<sub>C</sub>&#x00B7;<italic>C</italic>)+(&#x03B2;<sub>A</sub>+&#x03B2;<sub>AE</sub>&#x00B7;<italic>E</italic>+&#x03B2;<sub>AC</sub>&#x00B7;<italic>C</italic>)&#x00B7;<italic>A</italic>]&#x00D7;(1&#x2212;<italic>A</italic>)(3)</p><p>For a specific target error threshold <italic>T</italic>, <italic>A</italic><sub>threshold</sub> was computed as the smallest <italic>A</italic> &#x2208; [0, 1] at which error(<italic>A</italic>|<italic>E</italic>, <italic>C</italic>)&#x2264;<italic>T</italic>, obtained through a dense grid search over 1000 values of <italic>A</italic>. In total, 4 target error thresholds were evaluated, namely, 5%, 10%, 15%, and 20%. These values were not adopted from a single validated benchmark; rather, they constitute a graded series of progressively stricter error rate targets intended to span the range of error tolerances plausibly relevant to clinical decision support, from comparatively lenient targets appropriate to low-stakes supportive use to stringent targets appropriate to higher-stakes settings. Consistent with the exploratory aim of this study, the resulting <italic>A</italic><sub>threshold</sub> values should be interpreted as normative, hypothesis-generating estimates rather than as empirically validated error rate targets, and the specific threshold relevant to any given deployment will depend on the clinical context.</p></sec><sec id="s2-8"><title>Sensitivity Analysis</title><p>To assess model robustness, each literature-fixed coefficient (&#x03B2;<sub>E</sub>, &#x03B2;<sub>C</sub>, &#x03B2;<sub>AE</sub>, and &#x03B2;<sub>AC</sub>) was systematically perturbed by &#x00B1;20% from its baseline value. The calibrated coefficients (&#x03B2;<sub>&#x2080;</sub> and &#x03B2;<sub>A</sub>) were not perturbed because they were empirically determined rather than theoretically assumed. For each variant, the entire 27-cell simulation was rerun, and the resulting cell-level error rates were compared with baseline using Spearman rank correlation and maximum absolute error difference.</p></sec><sec id="s2-9"><title>Joint Human-AI Decision Model</title><p>The primary error specification (error=reliance&#x00D7;(1&#x2212;<italic>A</italic>)) attributes a clinical error solely to reliance on an incorrect AI recommendation and therefore represents the limiting case, in which the clinician contributes no independent accuracy. Human-AI teams do not always outperform either party alone and can even perform worse on decision-making tasks [<xref ref-type="bibr" rid="ref23">23</xref>,<xref ref-type="bibr" rid="ref24">24</xref>]&#x2014;for example, a randomized trial found that giving physicians access to an LLM did not significantly improve their diagnostic reasoning over conventional resources [<xref ref-type="bibr" rid="ref25">25</xref>]. In addition, users frequently overrely on AI by following incorrect recommendations [<xref ref-type="bibr" rid="ref13">13</xref>], and appropriate reliance requires accepting correct advice while rejecting incorrect advice [<xref ref-type="bibr" rid="ref26">26</xref>]. For these reasons, a joint human-AI decision model was additionally analyzed that represents both outcomes of an AI-assisted decision explicitly: when the clinician follows the AI, an error occurs if the AI is incorrect; and when the clinician overrides the AI, an error occurs if the clinician&#x2019;s own independent judgment is incorrect. Writing <italic>R</italic> for reliance (the probability of following the AI), <italic>A</italic> for AI accuracy, and <italic>p</italic><sub>c</sub> for the clinician&#x2019;s independent accuracy, the expected team error is error<sub>joint</sub>=<italic>R</italic>&#x00D7;(1&#x2212;<italic>A</italic>)+(1&#x2212;<italic>R</italic>)&#x00D7;(1&#x2212;<italic>p</italic><sub>c</sub>). The primary model is recovered exactly as the special case <italic>p</italic><sub>c</sub>&#x2192;1 (a clinician who is always correct when acting independently); this was verified to reproduce the reported <italic>A</italic><sub>threshold</sub> values (0.894 and 0.779 for the novice&#x00D7;high-complexity cell at the &#x003C;10% and &#x003C;20% targets). An asymmetric variant additionally distinguished passive deference to an incorrect recommendation from active correction of one by scaling the corrective term by a factor &#x03C1; &#x2208; [0, 1], where &#x03C1;=1 recovers the symmetric model, and smaller &#x03C1; represents clinicians who more often catch and overturn a wrong recommendation. <italic>A</italic><sub>threshold</sub> was recomputed under this model across clinician-accuracy levels <italic>p</italic><sub>c</sub> &#x2208; {0.6, 0.7, 0.8, 0.9, 1.0}.</p></sec><sec id="s2-10"><title>Extended Sensitivity Analysis</title><p>To probe robustness beyond the &#x00B1;20% perturbations, 3 additional analyses were performed. First, each literature-fixed coefficient was varied over a wide one-at-a-time range (&#x2212;100%, &#x2212;50%, +50%, and +100% of its baseline magnitude, together with a full sign reversal), and the 27-cell simulation was rerun for each variant. Second, a global Monte Carlo analysis simultaneously sampled all 4 literature-fixed coefficients from prior distributions across 50,000 draws under 2 priors: a theory-signed prior that preserved the assumed direction of each coefficient while allowing its magnitude to vary widely, and a fully sign-agnostic prior that additionally permitted every assumed sign to reverse. The resulting distribution of <italic>A</italic><sub>threshold</sub> values for the highest-risk (novice&#x00D7;high-complexity) and lowest-deference (experienced&#x00D7;low-complexity) cells was summarized by its median, 5th- to 95th-percentile interval, and the probability of exceeding an accuracy of 0.70. Third, to test whether the conclusions were an artifact of the linear reliance form, 2 bounded alternative structures&#x2014;a logistic (sigmoid) reliance function and a concave saturating function&#x2014;were fit to the same 9 calibration points, and <italic>A</italic><sub>threshold</sub> was recomputed from each. These analyses are reported in the Results section, summarized in <xref ref-type="supplementary-material" rid="app1">Multimedia Appendix 1</xref>, and provided in full in the repository.</p></sec></sec><sec id="s3" sec-type="results"><title>Results</title><sec id="s3-1"><title>Model Calibration</title><p>Weighted least-squares calibration of the reliance model against the 9 empirical data points yielded &#x03B2;<sub>&#x2080;</sub>=0.471 and &#x03B2;<sub>A</sub>=0.201. The calibrated model achieved an RMSE of 0.054 across the calibration targets (unweighted <italic>R</italic><sup>2</sup>=0.34, reflecting between-study heterogeneity in baseline reliance that the model does not attempt to capture). <xref ref-type="supplementary-material" rid="app2">Multimedia Appendix 2</xref> overlays the calibrated model predictions on the 9 empirical data points, stratified by the source study.</p><p>Study-level bootstrap resampling (2000 iterations) yielded a 95% CI for &#x03B2;<sub>A</sub> of 0.023&#x2010;0.234, which reflects the small number of independent calibration sources and the between-study heterogeneity in baseline reliance. Despite this width, the sign of &#x03B2;<sub>A</sub> was stable across all 2000 bootstrap iterations (<italic>P</italic>(&#x03B2;<sub>A</sub>&#x003E;0)&#x003E;.99), confirming that the monotonic positive relationship between AI accuracy and reliance is empirically robust. The leave-one-study-out cross-validation produced recalibrated &#x03B2;<sub>A</sub> values of 0.094 (excluding Lu and Yin [<xref ref-type="bibr" rid="ref15">15</xref>]), 0.246 (excluding experiment 3 by Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>]), and 0.199 (excluding experiment 1 phase 1 by Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>]). This finding indicates that the calibration depends most heavily on the study by Lu and Yin [<xref ref-type="bibr" rid="ref15">15</xref>] but does not collapse upon its removal. Per-study calibration yielded &#x03B2;<sub>A</sub> estimates of 0.023, 0.037, and 0.158, all positive: the joint calibration value of 0.201 represents the precision-weighted aggregate across these heterogeneous study-level estimates.</p><p>Two features of the calibration result warrant emphasis. First, the calibrated slope (&#x03B2;<sub>A</sub>=0.201) indicates that each 0.10-unit increase in AI accuracy corresponds to approximately a 0.02-unit increase in the predicted mean reliance. This is considerably weaker than the reliance-accuracy relationships assumed in some theoretically motivated frameworks and reflects the fact that human reliance on AI systems is sensitive to, but not dominated by, the accuracy of the system. Second, the 3 studies contributing calibration data have different baseline reliance levels (Lu and Yin [<xref ref-type="bibr" rid="ref15">15</xref>]: mean&#x2248;0.65; experiment 1 by Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>]: mean&#x2248;0.78; and experiment 3 by Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>]: mean&#x2248;0.82), a heterogeneity that the linear model averages across rather than resolves. Residuals were accordingly approximately symmetric around 0, but the RMSE was markedly larger than the SEs of the study means, an outcome we interpret as reflecting unmodeled study-level factors rather than misspecification of the functional form in the accuracy dimension.</p></sec><sec id="s3-2"><title>Predicted Reliance and Error Rates</title><p><xref ref-type="table" rid="table1">Table 1</xref> presents the predicted error rates across all 27 simulation conditions. The underlying predicted reliance values ranged from 0.421 (experienced clinicians on low-complexity tasks with <italic>A</italic>=0.5) to 0.938 (novice clinicians on high-complexity tasks with <italic>A</italic>=0.9), reflecting the combined influence of accuracy, experience, and complexity on the tendency to accept AI recommendations. Notably, the calibrated slope on <italic>A</italic> is sufficient to produce an approximately 0.08-unit increase in mean reliance between <italic>A</italic>=0.5 and <italic>A</italic>=0.9, but insufficient to generate an interior maximum of predicted error rate as a function of <italic>A</italic>.</p><table-wrap id="t1" position="float"><label>Table 1.</label><caption><p>Predicted error rates by simulation condition<sup><xref ref-type="table-fn" rid="table1fn1">a</xref></sup>.</p></caption><table id="table1" frame="hsides" rules="groups"><thead><tr><td align="left" valign="bottom">AI accuracy</td><td align="left" valign="bottom">Experience (years)</td><td align="left" valign="bottom">Low complexity, mean (SD)</td><td align="left" valign="bottom">Moderate complexity, mean (SD)</td><td align="left" valign="bottom">High complexity, mean (SD)</td></tr></thead><tbody><tr><td align="left" valign="top">0.5</td><td align="left" valign="top">0</td><td align="left" valign="top">0.29 (0.03)</td><td align="left" valign="top">0.35 (0.03)</td><td align="left" valign="top">0.41 (0.03)</td></tr><tr><td align="left" valign="top">0.5</td><td align="left" valign="top">2</td><td align="left" valign="top">0.27 (0.03)</td><td align="left" valign="top">0.33 (0.03)</td><td align="left" valign="top">0.40 (0.03)</td></tr><tr><td align="left" valign="top">0.5</td><td align="left" valign="top">10</td><td align="left" valign="top">0.21 (0.03)</td><td align="left" valign="top">0.27 (0.03)</td><td align="left" valign="top">0.34 (0.03)</td></tr><tr><td align="left" valign="top">0.7</td><td align="left" valign="top">0</td><td align="left" valign="top">0.18 (0.03)</td><td align="left" valign="top">0.22 (0.03)</td><td align="left" valign="top">0.26 (0.03)</td></tr><tr><td align="left" valign="top">0.7</td><td align="left" valign="top">2</td><td align="left" valign="top">0.18 (0.03)</td><td align="left" valign="top">0.22 (0.03)</td><td align="left" valign="top">0.26 (0.03)</td></tr><tr><td align="left" valign="top">0.7</td><td align="left" valign="top">10</td><td align="left" valign="top">0.14 (0.03)</td><td align="left" valign="top">0.19 (0.03)</td><td align="left" valign="top">0.23 (0.03)</td></tr><tr><td align="left" valign="top">0.9</td><td align="left" valign="top">0</td><td align="left" valign="top">0.07 (0.02)</td><td align="left" valign="top">0.08 (0.02)</td><td align="left" valign="top">0.09 (0.02)</td></tr><tr><td align="left" valign="top">0.9</td><td align="left" valign="top">2</td><td align="left" valign="top">0.06 (0.02)</td><td align="left" valign="top">0.08 (0.02)</td><td align="left" valign="top">0.09 (0.02)</td></tr><tr><td align="left" valign="top">0.9</td><td align="left" valign="top">10</td><td align="left" valign="top">0.05 (0.02)</td><td align="left" valign="top">0.07 (0.02)</td><td align="left" valign="top">0.08 (0.02)</td></tr></tbody></table><table-wrap-foot><fn id="table1fn1"><p><sup>a</sup>Values are mean (SD) of simulated error rates across 10,000 Monte Carlo trials per cell (seed=42).</p></fn></table-wrap-foot></table-wrap><p>The predicted error rates exhibited a clean 3-tier structure when grouped by AI accuracy level. At <italic>A</italic>=0.9, all 9 cells fell below a 10% error rate (range 0.054&#x2010;0.094). At <italic>A</italic>=0.7, the predicted error rates ranged from 0.145 to 0.265, spanning the 10%&#x2010;20% and 20%&#x2010;30% ranges, with all 3 novice cells exceeding 18% predicted error. At <italic>A</italic>=0.5, the predicted error rates ranged from 0.211 to 0.411, with all 6 moderate- and high-complexity cells exceeding 27% predicted error. For the cell most representative of the clinically high-risk scenario (novice clinicians performing high-complexity tasks), the predicted error rates were 0.411 at <italic>A</italic>=0.5, 0.265 at <italic>A</italic>=0.7, and 0.094 at <italic>A</italic>=0.9 (<xref ref-type="fig" rid="figure2">Figure 2A</xref>).</p><fig position="float" id="figure2"><label>Figure 2.</label><caption><p>Predicted error rates and accuracy thresholds for AI decision support in nursing. (A) Predicted mean error rate as a function of AI accuracy <italic>A</italic> at high task complexity (<italic>C</italic>=1.0) for 3 experience levels. The horizontal dotted lines mark 10%, 20%, and 30% target error rates. The shaded vertical band marks a reference range for the accuracy of contemporary general-purpose LLMs on complex clinical tasks (approximately 0.5&#x2010;0.7), based on published benchmarks [<xref ref-type="bibr" rid="ref18">18</xref>,<xref ref-type="bibr" rid="ref19">19</xref>]. (B) <italic>A</italic><sub>threshold</sub> curves plotting the required AI accuracy against the target error-rate threshold for 5 representative experience-by-complexity combinations. Solid and dashed lines show point estimates from the calibrated model. The shaded bands represent study-level bootstrap 95% CIs (2000 iterations). Horizontal reference lines mark a reference LLM accuracy level (0.70), plausible near-term frontier LLM accuracy (0.85), and demanding target accuracy (0.95). Note that <italic>A</italic><sub>threshold</sub> uncertainty is narrow at low target error rates (left side) and wider at higher target error rates (right side), reflecting the mathematical structure of error=reliance&#x00D7;(1&#x2212;<italic>A</italic>). LLM: large language model.</p></caption><graphic alt-version="no" mimetype="image" position="float" xlink:type="simple" xlink:href="nursing_v9i1e99590_fig02.png"/></fig><p>The <italic>A</italic>=0.7 simulation condition is closest to the accuracy levels currently reported for general-purpose LLMs on complex clinical tasks [<xref ref-type="bibr" rid="ref18">18</xref>,<xref ref-type="bibr" rid="ref19">19</xref>], in which the predicted error rates for novice clinicians range from 0.184 (low-complexity tasks) to 0.265 (high-complexity tasks). At high task complexity, the condition most relevant to ambiguous clinical presentations, the predicted error rate of 0.265 implies that approximately 1 in 4 AI-supported clinical decisions would be incorrect when the system is used by a novice clinician at this accuracy level.</p></sec><sec id="s3-3"><title><italic>A</italic><sub>threshold</sub> Analysis</title><p><xref ref-type="table" rid="table2">Table 2</xref> presents <italic>A</italic><sub>threshold</sub> values, the minimum AI accuracy required to achieve a specified error rate, for each of the 9 experience-by-complexity combinations at 4 target error thresholds (5%, 10%, 15%, and 20%). For novice clinicians performing high-complexity tasks, the minimum AI accuracy values required to achieve error rates &#x003C;10%, &#x003C;15%, and &#x003C;20% were 0.894, 0.838, and 0.779, respectively. Corresponding values for experienced clinicians performing low-complexity tasks were 0.806, 0.686, and 0.539, respectively. These differences highlight the interaction between clinician experience, task complexity, and the minimum accuracy requirement for clinical use. Because the reliance model is calibrated on nonnursing behavioral data, these numerical thresholds should be read as illustrative outputs of the model rather than as nursing-derived empirical standards; their value lies in the relative ordering across user-and-task profiles rather than in the specific accuracy figures.</p><table-wrap id="t2" position="float"><label>Table 2.</label><caption><p><italic>A</italic><sub>threshold</sub> values&#x2014;minimum AI accuracy required to achieve specified error rate targets<sup><xref ref-type="table-fn" rid="table2fn1">a</xref></sup>.</p></caption><table id="table2" frame="hsides" rules="groups"><thead><tr><td align="left" valign="bottom">Experience and complexity</td><td align="left" valign="bottom">Error&#x003C;5%</td><td align="left" valign="bottom">Error&#x003C;10%</td><td align="left" valign="bottom">Error&#x003C;15%</td><td align="left" valign="bottom">Error&#x003C;20%</td></tr></thead><tbody><tr><td align="char" char="." valign="top" colspan="5">0 years</td></tr><tr><td align="left" valign="top"><named-content content-type="indent">&#x00A0;&#x00A0;&#x00A0;&#x00A0;</named-content>Low</td><td align="left" valign="top">0.925</td><td align="left" valign="top">0.845</td><td align="left" valign="top">0.759</td><td align="left" valign="top">0.670</td></tr><tr><td align="left" valign="top"><named-content content-type="indent">&#x00A0;&#x00A0;&#x00A0;&#x00A0;</named-content>Moderate</td><td align="left" valign="top">0.939</td><td align="left" valign="top">0.874</td><td align="left" valign="top">0.807</td><td align="left" valign="top">0.736</td></tr><tr><td align="left" valign="top"><named-content content-type="indent">&#x00A0;&#x00A0;&#x00A0;&#x00A0;</named-content>High</td><td align="left" valign="top">0.949</td><td align="left" valign="top">0.894</td><td align="left" valign="top">0.838</td><td align="left" valign="top">0.779</td></tr><tr><td align="char" char="." valign="top" colspan="5">2 years</td></tr><tr><td align="left" valign="top"><named-content content-type="indent">&#x00A0;&#x00A0;&#x00A0;&#x00A0;</named-content>Low</td><td align="left" valign="top">0.922</td><td align="left" valign="top">0.839</td><td align="left" valign="top">0.749</td><td align="left" valign="top">0.653</td></tr><tr><td align="left" valign="top"><named-content content-type="indent">&#x00A0;&#x00A0;&#x00A0;&#x00A0;</named-content>Moderate</td><td align="left" valign="top">0.937</td><td align="left" valign="top">0.870</td><td align="left" valign="top">0.800</td><td align="left" valign="top">0.726</td></tr><tr><td align="left" valign="top"><named-content content-type="indent">&#x00A0;&#x00A0;&#x00A0;&#x00A0;</named-content>High</td><td align="left" valign="top">0.947</td><td align="left" valign="top">0.891</td><td align="left" valign="top">0.834</td><td align="left" valign="top">0.773</td></tr><tr><td align="char" char="." valign="top" colspan="5">10 years</td></tr><tr><td align="left" valign="top"><named-content content-type="indent">&#x00A0;&#x00A0;&#x00A0;&#x00A0;</named-content>Low</td><td align="left" valign="top">0.909</td><td align="left" valign="top">0.806</td><td align="left" valign="top">0.686</td><td align="left" valign="top">0.539</td></tr><tr><td align="left" valign="top"><named-content content-type="indent">&#x00A0;&#x00A0;&#x00A0;&#x00A0;</named-content>Moderate</td><td align="left" valign="top">0.929</td><td align="left" valign="top">0.851</td><td align="left" valign="top">0.766</td><td align="left" valign="top">0.670</td></tr><tr><td align="left" valign="top"><named-content content-type="indent">&#x00A0;&#x00A0;&#x00A0;&#x00A0;</named-content>High</td><td align="left" valign="top">0.942</td><td align="left" valign="top">0.879</td><td align="left" valign="top">0.812</td><td align="left" valign="top">0.740</td></tr></tbody></table><table-wrap-foot><fn id="table2fn1"><p><sup>a</sup><italic>A</italic><sub>threshold</sub> is the smallest AI accuracy value <italic>A</italic> &#x2208; [0, 1] at which the calibrated model&#x2019;s predicted mean error rate falls below the specified error rate target, computed analytically from the closed-form quadratic expression derived from the linear reliance specification. Values are rounded to 3 decimal places and were computed from the full-precision calibrated coefficients; exact coefficient values are available in the repository.</p></fn></table-wrap-foot></table-wrap><p>Bootstrap propagation of calibration uncertainty into <italic>A</italic><sub>threshold</sub> values revealed an important asymmetry: although &#x03B2;<sub>A</sub> has a wide bootstrap 95% CI, the <italic>A</italic><sub>threshold</sub> values are tightly bounded. For the novice&#x00D7;high-complexity combination, the bootstrap 95% CIs for the <italic>A</italic><sub>threshold</sub> to achieve error rates of 10%, 15%, and 20% were 0.877&#x2010;0.898, 0.814&#x2010;0.845, and 0.750&#x2010;0.792, respectively. This tight propagation reflects the mathematical structure of error=reliance&#x00D7;(1&#x2212;<italic>A</italic>): at low target error rates, the (1&#x2212;<italic>A</italic>) factor must be small, which constrains <italic>A</italic> to lie close to 1 regardless of the exact value of &#x03B2;<sub>A</sub> within its CI. Correspondingly, across the accuracy range currently reported for general-purpose LLMs on complex clinical tasks [<xref ref-type="bibr" rid="ref18">18</xref>,<xref ref-type="bibr" rid="ref19">19</xref>], the predicted error rate for novice clinicians on high-complexity tasks ranges from approximately 0.26 at the upper end of that range to approximately 0.41 at the lower end. The conclusion that AI systems operating in this accuracy range yield error rates well above stringent targets for this user-and-task combination is robust across the bootstrap distribution of the calibrated coefficients.</p><p><xref ref-type="fig" rid="figure2">Figure 2B</xref> presents the <italic>A</italic><sub>threshold</sub> function graphically, plotting the required AI accuracy against the target error-rate threshold for 5 representative experience-by-complexity combinations. Horizontal reference lines mark a reference LLM accuracy level (0.70), plausible near-term frontier LLM accuracy (0.85), and demanding target accuracy (0.95). At a reference accuracy level of 0.70, a total of 4 of the 9 cells achieve predicted error rates &#x003C;20%: the 3 low-complexity cells (ranging from 14.5% for experienced clinicians to 18.4% for novices) and the experienced&#x00D7;moderate-complexity cell (18.5%). The remaining 5 cells&#x2014;comprising novice and 2-year-experience clinicians on moderate- or high-complexity tasks, plus experienced clinicians on high-complexity tasks&#x2014;have predicted error rates ranging from 21.6% to 26.5%, consistently exceeding a 20% error rate.</p></sec><sec id="s3-4"><title>Sensitivity Analysis</title><p>When each of the 4 literature-fixed coefficients (&#x03B2;<sub>E</sub>, &#x03B2;<sub>C</sub>, &#x03B2;<sub>AE</sub>, and &#x03B2;<sub>AC</sub>) was varied by &#x00B1;20% from its baseline value, the overall error ranking of the 27 cells was nearly preserved in every perturbation. Across all 8 perturbations (4 coefficients&#x00D7;2 directions), the Spearman rank correlation between the perturbed and baseline cell-level error rankings ranged from 0.991 to 1.000, with a mean of 0.998 (SD 0.003). The maximum absolute change in any cell&#x2019;s predicted error rate across all perturbations was 0.020, occurring for the &#x03B2;<sub>E</sub>+20% perturbation in the experienced&#x00D7;moderate-complexity&#x00D7;<italic>A</italic>=0.5 cell (baseline 0.273; perturbed 0.253). These results indicate that the principal conclusions of the simulation&#x2014;the tiered error-rate structure of the 27 cells and the <italic>A</italic><sub>threshold</sub> values reported in <xref ref-type="table" rid="table2">Table 2</xref>&#x2014;are robust to moderate perturbation of the literature-fixed coefficients.</p></sec><sec id="s3-5"><title>Joint Decision Model and Extended Robustness</title><p>Under the joint human-AI decision model, the principal conclusion for the highest-risk condition was unchanged. For novice clinicians performing high-complexity tasks, the minimum AI accuracy required to keep predicted team error below 10% moved only from 0.894 at <italic>p</italic><sub>c</sub>&#x2192;1 to 0.917 at <italic>p</italic><sub>c</sub>=0.6, and the threshold for the&#x003C;20% target moved from 0.779 to 0.818 across the same range; because reliance is high in this cell, the override path contributes little, and clinician accuracy has limited leverage. A qualitatively different pattern emerged in the lowest-deference cell (experienced clinicians on low-complexity tasks), where team error acquired an irreducible floor equal to (1&#x2212;<italic>R</italic>(<italic>A</italic>=1))&#x00D7;(1&#x2212;<italic>p</italic><sub>c</sub>). This floor grew from 0% at <italic>p</italic><sub>c</sub>&#x2192;1 to 8.6% at <italic>p</italic><sub>c</sub>=0.8 and 17.1% at <italic>p</italic><sub>c</sub>=0.6, so that the&#x003C;10% target became unreachable by any AI accuracy once <italic>p</italic><sub>c</sub> fell to 0.7 or below. This structural feature&#x2014;that improving AI accuracy alone cannot drive team error below the clinician&#x2019;s own error contribution&#x2014;indicates that a single accuracy number is insufficient and that the human contribution to the joint decision must be represented. Under the asymmetric variant, allowing clinicians to correct a wrong recommendation (&#x03C1;&#x003C;1) lowered the required accuracy substantially (for the novice&#x00D7;high-complexity cell at the &#x003C;20% target, from 0.779 at &#x03C1;=1 to 0.673 at &#x03C1;=0.7 and 0.516 at &#x03C1;=0.5), indicating that the primary model is conservative with respect to correction.</p><p>The extended sensitivity analyses reinforced the &#x00B1;20% results. Under wide one-at-a-time perturbation (up to &#x00B1;100% and sign reversal of each literature-fixed coefficient), the novice&#x00D7;high-complexity threshold was mathematically invariant to &#x03B2;<sub>E</sub> and &#x03B2;<sub>A</sub> (both multiply the experience term, which is 0 in that cell) and varied only with the complexity-related coefficients; even under a full sign reversal of &#x03B2;<sub>C</sub>, the required accuracy for the &#x003C;10% target remained between 0.805 and 0.900, and the rank ordering of the 27 cells was preserved (Spearman &#x03C1;&#x2265;0.77 across all perturbations, and &#x03C1;&#x003E;0.95 for most). In the global Monte Carlo analysis (50,000 draws per prior), the theory-signed prior yielded a novice&#x00D7;high-complexity <italic>A</italic><sub>threshold</sub> with median 0.894 (5th-95th percentile 0.863&#x2010;0.900) for the &#x003C;10% target and 0.780 (5th-95th percentile 0.709&#x2010;0.800) for the &#x003C;20% target, with the required accuracy exceeding 0.70 in 100% and 97.4% of draws, respectively. Even under the fully sign-agnostic prior, in which every assumed coefficient sign was permitted to reverse, the &#x003C;10% threshold for this cell retained a median of 0.844 and exceeded 0.70 in 84.4% of the sampled parameter space; values in the low tail arose only under sign combinations contrary to the theoretical expectations. Finally, both bounded alternative reliance structures fit the 9 calibration points essentially as well as the linear form (RMSE 0.0542 for the logistic and concave-saturating forms vs 0.054 for the linear form) and produced near-identical thresholds for the novice&#x00D7;high-complexity cell (0.900 and 0.794&#x2010;0.795 for the &#x003C;10% and &#x003C;20% targets, respectively), indicating that the conclusions are not an artifact of the linear specification. <xref ref-type="supplementary-material" rid="app1">Multimedia Appendix 1</xref> summarizes the joint-model and global-sensitivity results.</p></sec></sec><sec id="s4" sec-type="discussion"><title>Discussion</title><sec id="s4-1"><title>Principal Findings</title><p>In this study, an empirically calibrated simulation model was developed to quantify the AI accuracy required to achieve specified error rate targets in nursing decision-making across combinations of clinician experience and task complexity. Three principal findings warrant emphasis. We emphasize at the outset that the contribution of this work is a reusable, model-agnostic theoretical framework rather than an evaluation of any particular AI system. Its value lies in the mapping it defines from AI accuracy, clinician experience, and task complexity to a predicted error rate and a corresponding minimum-accuracy requirement&#x2014;a mapping that does not depend on which model is current or on the specific accuracy figure entered into it.</p><p>First, the moderate accuracy levels currently reported for general-purpose LLMs on complex clinical tasks [<xref ref-type="bibr" rid="ref18">18</xref>,<xref ref-type="bibr" rid="ref19">19</xref>] fall within an operating zone where the predicted error rates exceed stringent targets for novice clinicians performing complex tasks. At the upper end of this reported range (approximately 0.70), the simulation predicts an error rate of approximately 0.26 in the novice&#x00D7;high-complexity cell&#x2014;rising to roughly 0.41 at the lower end of the range&#x2014;substantially exceeding the stringent error targets relevant to autonomous clinical decision support in high-risk clinical scenarios.</p><p>Second, the calibrated reliance model&#x2014;grounded in 9 empirical data points from 3 independent randomized experiments (total N=3502)&#x2014;indicates that human reliance on AI recommendations rises monotonically with AI accuracy but at a modest rate (&#x03B2;<sub>A</sub>=0.20). This slope is considerably shallower than the reliance-accuracy relationships assumed in some theoretical frameworks and reflects the cumulative evidence that users calibrate their reliance in response to AI accuracy, but do so imperfectly and with substantial residual reliance even on low-accuracy systems. Because reliance does not rise steeply enough to offset the declining AI failure probability at high accuracy levels, the predicted clinical error rates in this model decrease monotonically as AI accuracy increases, without an interior maximum. The clinical implication is that no &#x201C;dangerous middle&#x201D; region exists in which error rates peak; the clinical concern, rather, is the absolute magnitude of error rates across a broad low-to-moderate accuracy range.</p><p>Third, <italic>A</italic><sub>threshold</sub> values, the minimum AI accuracy required to achieve a specified error rate target, exhibit a pronounced dependence on clinician experience and task complexity. For the highest-risk combination (novice clinicians performing high-complexity tasks), achieving an error rate &#x003C;10% requires AI accuracy of at least 0.89, and achieving an error rate &#x003C;20% requires accuracy of at least 0.78. For experienced clinicians performing low-complexity tasks, the corresponding thresholds are 0.81 and 0.54. These differences, spanning up to 24 percentage points of the required accuracy across user-and-task profiles (at the &#x003C;20% error target), indicate that a single accuracy benchmark cannot meaningfully govern clinical use decisions; the threshold depends on the user-and-task context.</p></sec><sec id="s4-2"><title>Comparison With Prior Work</title><p>The predictions of the proposed model converge with findings from 2 independent lines of empirical research on human-AI interaction.</p><p>First, in the AI-assisted decision-making literature, the calibrated model is grounded directly in the reliance data from Lu and Yin [<xref ref-type="bibr" rid="ref15">15</xref>] and Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>], representing 3 randomized experiments with a combined sample of 3502 participants. The linear form of the model aligns with the empirically observed monotonic relationship between AI accuracy and human reliance, and the absence of an interior reliance maximum in the range <italic>A</italic> &#x2208; [0.5, 1.0], observed in both studies, directly shaped the structural form of the reliance equation. This empirical basis distinguishes the present framework from earlier theoretical treatments that posited nonlinear reliance patterns without supporting behavioral data.</p><p>Second, in the clinical decision support literature, K&#x00FC;cking et al [<xref ref-type="bibr" rid="ref16">16</xref>] reported that AI recommendation correctness exerted a strong bidirectional influence on the diagnostic performance of 223 physicians and nurses (OR 10.0; <italic>P</italic>&#x003C;.001 for correct AI recommendations; reciprocal declines for incorrect recommendations). They also found that longer work experience was associated with higher diagnostic accuracy (OR 1.89), and formal qualifications further improved accuracy (OR 1.40). These clinical expertise effects are directionally consistent with the negative experience coefficient (&#x03B2;<sub>E</sub>=&#x2212;0.20) specified a priori in the present model. In a related study by the same group [<xref ref-type="bibr" rid="ref27">27</xref>], the overall human diagnostic accuracy averaged 79.3% (up to 85% in the formally qualified subgroup); by comparison, general-purpose LLMs have been reported to achieve lower accuracy on complex clinical tasks [<xref ref-type="bibr" rid="ref18">18</xref>,<xref ref-type="bibr" rid="ref19">19</xref>], situating the operating range explored in the present simulation within a clinically meaningful band relative to human performance.</p><p>The introduced <italic>A</italic><sub>threshold</sub> framework is best understood in relation to 3 established frameworks for human-automation interaction. Lee and See [<xref ref-type="bibr" rid="ref10">10</xref>] proposed a process-oriented framework in which trust calibration determines appropriate reliance but did not provide quantitative guidance for translating system accuracy into use case decisions. Parasuraman and Riley [<xref ref-type="bibr" rid="ref9">9</xref>] articulated a use/misuse/disuse/abuse taxonomy that classifies human-automation interaction failure modes but treats accuracy as an input among many rather than as a use-determining quantity. More recently, the literature on trustworthy AI [<xref ref-type="bibr" rid="ref6">6</xref>] has emphasized explainability, fairness, and accountability as design principles, but has generally not specified accuracy thresholds for clinical use.</p><p>The present framework differs from prior approaches in 3 ways. First, it converts a single quantitative input&#x2014;AI accuracy on the intended task&#x2014;into a deployment decision using an empirically calibrated function. Second, it explicitly adjusts the deployment threshold based on user experience and task complexity, recognizing that no universal accuracy cutoff is appropriate across different use cases. Third, the graphical guidance provided in <xref ref-type="fig" rid="figure2">Figure 2B</xref> allows clinicians and administrators to read off the minimum required accuracy for a specific error tolerance and user-task profile without requiring statistical expertise. These properties make the <italic>A</italic><sub>threshold</sub> framework complementary to, rather than competitive with, existing process- and design-oriented frameworks: <italic>A</italic><sub>threshold</sub> provides a quantitative gating criterion that can be applied alongside trust calibration and trustworthy AI principles in deployment review processes.</p><p>Taken together, the present framework is consistent with both the quantitative reliance patterns observed in general-purpose AI-assisted decision-making and the qualitative expertise effects observed in clinical decision support. The principal gap in the empirical foundation&#x2014;behavioral reliance data collected specifically in nursing populations performing clinical decision-making tasks across multiple AI accuracy levels&#x2014;remains a priority for future research.</p></sec><sec id="s4-3"><title>Theoretical Integration</title><p>The finding that <italic>A</italic><sub>threshold</sub> varies strongly with task complexity can be interpreted through the lens of dual-process theory [<xref ref-type="bibr" rid="ref28">28</xref>]. At low task complexity, system 2 (analytical) processing may remain engaged even when AI accuracy is moderate, because the clinician retains sufficient cognitive resources to evaluate the AI recommendation independently. At high task complexity, by contrast, system 1 (automatic) processing may predominate, as cognitive resources are preferentially allocated to the task itself rather than to critical evaluation of the AI recommendation. This asymmetry provides a theoretical rationale for the model&#x2019;s larger complexity-by-accuracy interaction coefficient and for the model-implied result that novice clinicians, who have not yet developed efficient system 2 strategies for clinical reasoning [<xref ref-type="bibr" rid="ref21">21</xref>,<xref ref-type="bibr" rid="ref22">22</xref>], face the most stringent AI accuracy requirements.</p><p>Within the framework of Lee and See [<xref ref-type="bibr" rid="ref10">10</xref>], the model implies that appropriate reliance is unlikely to be fully achieved through user education or interface design alone when AI accuracy falls substantially below the <italic>A</italic><sub>threshold</sub> for the relevant user-and-task combination. In such contexts, the residual error rate is a function of system accuracy rather than of user awareness, and the appropriate intervention is either to restrict AI use to user-and-task combinations for which the <italic>A</italic><sub>threshold</sub> is met or to improve AI accuracy itself.</p><p>A further implication concerns the durability of the framework, as language models continue to improve. Because the <italic>A</italic><sub>threshold</sub> framework maps any given AI accuracy onto a predicted error rate and a corresponding minimum-accuracy requirement, its utility does not depend on the accuracy level of any particular contemporary model. As model performance advances, the same framework can be reapplied to locate an improved system within the accuracy-error landscape and to re-evaluate which user-and-task combinations it can safely support. The specific operating range attributed here to current general-purpose LLMs is thus a movable reference point within a framework that is itself invariant to such change.</p></sec><sec id="s4-4"><title>Clinical Implications</title><p>The following implications are exploratory and contingent on behavioral validation in nursing settings; they describe how the <italic>A</italic><sub>threshold</sub> framework might inform practice rather than offering validated recommendations and should be interpreted in conjunction with the limitations. In particular, a system&#x2019;s meeting a given <italic>A</italic><sub>threshold</sub> on a benchmark should be treated as a necessary but not a sufficient condition for nursing use, and the framework must not be interpreted as authorization to deploy a system without behavioral validation in the relevant nursing context.</p><p>First, the minimum AI accuracy required for clinical use should be specified as a function of the intended user and task, not as a single benchmark. Although an AI system operating at the moderate accuracy levels typical of current general-purpose LLMs on complex clinical tasks [<xref ref-type="bibr" rid="ref18">18</xref>,<xref ref-type="bibr" rid="ref19">19</xref>] may be suitable for use by experienced clinicians in low-complexity tasks, it is predicted to produce error rates &#x003E;25% when used by novice clinicians in high-complexity decision-making. Therefore, clinical use decisions should be based on the specific <italic>A</italic><sub>threshold</sub> value that corresponds to the intended user population and task profile, rather than on the raw AI accuracy value alone. Moreover, the tolerable error rate is itself context-dependent&#x2014;shaped by the clinical setting, the severity of potential harm, the degree of human supervision, and whether the system is used in a supportive or an autonomous capacity&#x2014;so a target that is reasonable for an educational support tool may be inadequate for high-risk acute care.</p><p>Second, as AI is increasingly integrated across nursing practice and curricular reform has been called for to prepare nurses to work safely alongside these technologies [<xref ref-type="bibr" rid="ref29">29</xref>], nursing education programs should incorporate explicit instruction on the relationship among AI accuracy, user experience, and clinical error rates. Students should understand that AI systems in the low-to-moderate accuracy range typical of current general-purpose LLMs may be inadequate for autonomous use in high-complexity clinical decisions. In these settings, AI should augment, not replace, clinical judgment. The <italic>A</italic><sub>threshold</sub> framework offers a practical decision tool for this educational purpose.</p><p>Third, clinical workflows could adopt differentiated deployment protocols informed by the <italic>A</italic><sub>threshold</sub> framework. When the current AI accuracy for user-and-task combinations falls below the relevant <italic>A</italic><sub>threshold</sub>, enhanced monitoring safeguards, such as mandatory verification steps, second-clinician review, or confidence-based routing, should be applied preferentially. These measures provide a practical intermediate solution while AI accuracy continues to increase toward levels suitable for autonomous use.</p><p>Fourth, regulatory and institutional review processes for clinical AI decision support could be informed by encouraging deployment proposals to specify the intended user population and task profile and to consider the corresponding <italic>A</italic><sub>threshold</sub>. Current regulatory frameworks do not typically request this level of user-and-task-specific justification; the <italic>A</italic><sub>threshold</sub> framework is offered as a candidate quantitative input to such review, pending behavioral validation, rather than as a validated regulatory standard. This need is underscored by recent calls for explicit regulatory oversight of LLMs used in health care [<xref ref-type="bibr" rid="ref30">30</xref>,<xref ref-type="bibr" rid="ref31">31</xref>].</p></sec><sec id="s4-5"><title>Limitations</title><p>Several limitations merit consideration.</p><p>First, although the accuracy-related coefficients of the reliance model were empirically calibrated against published behavioral data, the framework itself remains a theoretical construct, and its predictions require direct behavioral validation. Calibration confirms that the monotonic reliance slope is consistent with empirical observations in AI-assisted decision-making, but it does not establish that the noncalibrated coefficients (&#x03B2;<sub>E</sub>, &#x03B2;<sub>C</sub>, &#x03B2;<sub>AE</sub>, and &#x03B2;<sub>AC</sub>) accurately reflect nurses performing clinical tasks. The same caution applies to the joint decision model introduced in response to review: the clinician&#x2019;s independent accuracy (<italic>p</italic><sub>c</sub>) and the correction factor (&#x03C1;) are illustrative parameters rather than calibrated quantities, so the joint-model results should be read as a structured sensitivity analysis over a plausible range of clinician behavior&#x2014;showing which conclusions are robust to, and which depend on, the human contribution to the decision&#x2014;rather than as point predictions for any specific clinician accuracy. The framework should therefore be interpreted as a decision-support and hypothesis-generating tool rather than as a validated predictive model for individual clinical use.</p><p>Second, the accuracy levels used to contextualize the simulation&#x2019;s operating range were drawn from published benchmarks of general-purpose LLMs on complex clinical tasks rather than from a nursing-specific evaluation. Reported LLM accuracy varies substantially across models, prompting strategies, task types, and clinical domains, and the accuracy of frontier LLMs on complex clinical tasks remains an active area of research. The operating range examined here should therefore be regarded as an illustrative band informed by the current literature rather than as a definitive estimate of any specific system&#x2019;s clinical performance.</p><p>Third, calibration data were drawn from 3 studies that differed systematically in baseline reliance levels (Lu and Yin [<xref ref-type="bibr" rid="ref15">15</xref>]: ~0.65; experiment 1 by Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>]: ~0.78; and experiment 3 by Yin et al [<xref ref-type="bibr" rid="ref14">14</xref>]: ~0.82). The linear model averages across this heterogeneity, producing an RMSE (0.054) larger than the SEs of individual study means. Therefore, the calibrated slope (&#x03B2;<sub>A</sub>=0.201) should be interpreted as a central tendency across heterogeneous contexts rather than a precise parameter applicable to any specific setting. Part of this heterogeneity may reflect methodological differences in how AI accuracy was manipulated (stated, observed, or designed), which the calibration pooled as proxies for the underlying <italic>A</italic> variable. The heterogeneity is informative and motivates future research parameterizing contextual factors beyond AI accuracy. With only 3 independent calibration studies, the study-level bootstrap has a limited sample space, so the reported CIs should be viewed as lower bounds on the true uncertainty. Consistent with this, the <italic>A</italic><sub>threshold</sub> values are reported to 3 decimal places to reflect the numerical precision of the simulation, not the empirical precision of the underlying parameter, which is bounded by a wide CI whose lower limit approaches 0. The thresholds are therefore best interpreted in relative and ordinal terms&#x2014;comparing the accuracy demands of different user-and-task profiles&#x2014;rather than as exact accuracy targets to be applied at the reported number of significant figures.</p><p>Fourth, the calibration data were derived from studies in which crowdworkers predicted speed-dating outcomes rather than clinicians making clinical decisions. Although this is the most empirically rich reliance-by-accuracy source currently available, the clinical context gap is substantial. Clinicians may exhibit baseline reliance, accuracy sensitivities, and experience effects that differ from those of crowdworkers, and the task complexity assigned to speed-dating prediction (<italic>C</italic>=0.5) reflects a qualitative judgment rather than direct measurement. Even so, the directional structure of the model is consistent with recent experimental evidence from clinical populations that include nurses. In a web-based experiment with 223 dermatologists, reliance on AI was stronger for correct than for incorrect advice and decreased with greater medical experience [<xref ref-type="bibr" rid="ref32">32</xref>]. More directly, in a wound-maceration study, 223 physicians and nurses made decisions that were sensitive to whether the AI recommendation was correct or incorrect [<xref ref-type="bibr" rid="ref16">16</xref>], and a related report measuring the rate of agreement with incorrect recommendations found that higher diagnostic performance and certified wound-care training were associated with lower agreement [<xref ref-type="bibr" rid="ref33">33</xref>]. These patterns match the positive accuracy-reliance slope and the experience-based attenuation encoded in the model, supporting cautious extrapolation of its directional structure to clinical settings. Nursing-specific calibration data obtained across multiple AI accuracy levels would substantially strengthen the empirical foundation of the framework.</p><p>Fifth, the linear reliance specification was chosen for parsimony and tractability, since 9 aggregate calibration points are insufficient to identify more elaborate functional forms. Alternative specifications (logistic, power-law, and segmented-linear) might provide superior fit with richer data. The linear form should therefore be interpreted as the simplest specification consistent with the available data, not as a claim that the underlying reliance-accuracy relationship is truly linear.</p><p>Sixth, the model does not incorporate contextual factors&#x2014;such as time pressure, team dynamics, organizational accountability, and environmental stressors&#x2014;that are known to influence clinical decision-making. These factors likely contribute substantially to real-world reliance and to the observed between-study heterogeneity. Future investigations should explicitly parameterize the most influential contextual factors, which will require more calibration data than are currently available.</p><p>Seventh, the reliance and error expressions are clipped to [0, 1] to preserve their probabilistic interpretation. Within the reported 27-cell grid, clipping affected &#x003C;1% of simulated values and did not materially alter the reported results. At extreme accuracy levels outside the reported grid, clipping becomes more frequent, and the linear reliance expression should not be extrapolated beyond its reported domain.</p></sec><sec id="s4-6"><title>Conclusions</title><p>This study provides an empirically calibrated simulation framework for quantifying the AI accuracy required to achieve specified error rate targets in nursing decision-making. Calibration against 9 empirical data points derived from 3 independent randomized experiments (total N=3502) indicates that human reliance on AI rises gradually but modestly with AI accuracy. For the highest-risk combination&#x2014;novice clinicians performing high-complexity tasks&#x2014;the predicted clinical error rates remain above a stringent target (eg, &#x003C;10%) when AI accuracy is below approximately 0.89. The moderate accuracy levels currently reported for general-purpose LLMs on complex clinical tasks [<xref ref-type="bibr" rid="ref18">18</xref>,<xref ref-type="bibr" rid="ref19">19</xref>] fall well below this threshold, suggesting that, on present evidence, such systems are not yet adequate for autonomous use by novice nurses on complex clinical decisions and should serve as an adjunct to, rather than a replacement for, clinical judgment. The <italic>A</italic><sub>threshold</sub> framework introduced here provides a decision-theoretic tool for specifying minimum accuracy requirements as a function of clinician experience and task complexity, and it can inform AI deployment, nursing education, and regulatory policy. Direct behavioral validation in nursing populations remains an essential next step.</p></sec></sec></body><back><ack><p>The author thanks Lu and Yin and Yin et al for making their experimental data and figures publicly available, which made empirical calibration possible. During the preparation of this work, the author used Claude (Anthropic) as a writing assistance tool to draft and revise sections of the manuscript in English, assist in the analysis of publicly available empirical reliance data, and generate Python code for model calibration and Monte Carlo simulation. After using this tool, the author reviewed and edited the content as needed and took full responsibility for the content of the publication.</p></ack><notes><sec><title>Funding</title><p>The author declared no financial support was received for this work.</p></sec><sec><title>Data Availability</title><p>The data and code that support the findings of this study are openly available in the nursing-AI-error-thresholds repository on GitHub under the MIT License [<xref ref-type="bibr" rid="ref34">34</xref>]. The complete reproducibility package includes all calibration scripts, the bootstrap procedure (2000 study-level iterations), the 27-cell factorial simulation, the joint decision model, and the extended (wide, global, and alternative-structure) sensitivity analyses, together with vector-format source files for every figure reported in this paper. Rerunning the figure-generation script from a clean Python &#x2265;3.10 environment reproduces the calibrated coefficients (&#x03B2;<sub>0</sub>=0.4708; &#x03B2;<sub>A</sub>=0.2010), the bootstrap 95% CI for &#x03B2;<sub>A</sub> (0.023-0.234), and all reported numerical results to at least 3 decimal places across platforms. The pseudorandom number generator was initialized with seed 42 using the NumPy default_rng interface, and the 27-cell factorial simulation used 10,000 Monte Carlo trials per cell. Analyses were conducted in Python (version 3.11) with NumPy (version 1.26) and SciPy (version 1.13). A preprint of this manuscript is available on SSRN [<xref ref-type="bibr" rid="ref35">35</xref>].</p></sec></notes><fn-group><fn fn-type="con"><p>HT is the sole author and contributed the following roles according to the CRediT taxonomy: conceptualization, methodology, software, formal analysis, investigation, data curation, writing&#x2014;original draft, writing&#x2014;review and editing, and visualization. The author had full access to all data and analyses and takes full responsibility for the integrity of the work and the decision to submit the manuscript for publication.</p></fn><fn fn-type="conflict"><p>None declared.</p></fn></fn-group><glossary><title>Abbreviations</title><def-list><def-item><term id="abb1">LLM</term><def><p>large language model</p></def></def-item><def-item><term id="abb2">OR</term><def><p>odds ratio</p></def></def-item><def-item><term id="abb3">RMSE</term><def><p>root-mean-square error</p></def></def-item></def-list></glossary><ref-list><title>References</title><ref id="ref1"><label>1</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Topol</surname><given-names>EJ</given-names> </name></person-group><article-title>High-performance medicine: the convergence of human and artificial intelligence</article-title><source>Nat Med</source><year>2019</year><month>01</month><volume>25</volume><issue>1</issue><fpage>44</fpage><lpage>56</lpage><pub-id pub-id-type="doi">10.1038/s41591-018-0300-7</pub-id><pub-id pub-id-type="medline">30617339</pub-id></nlm-citation></ref><ref id="ref2"><label>2</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Blease</surname><given-names>C</given-names> </name><name name-style="western"><surname>Kaptchuk</surname><given-names>TJ</given-names> </name><name name-style="western"><surname>Bernstein</surname><given-names>MH</given-names> </name><name name-style="western"><surname>Mandl</surname><given-names>KD</given-names> </name><name name-style="western"><surname>Halamka</surname><given-names>JD</given-names> </name><name name-style="western"><surname>DesRoches</surname><given-names>CM</given-names> </name></person-group><article-title>Artificial intelligence and the future of primary care: exploratory qualitative study of UK general practitioners&#x2019; views</article-title><source>J Med Internet Res</source><year>2019</year><month>03</month><day>20</day><volume>21</volume><issue>3</issue><fpage>e12802</fpage><pub-id pub-id-type="doi">10.2196/12802</pub-id><pub-id pub-id-type="medline">30892270</pub-id></nlm-citation></ref><ref id="ref3"><label>3</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Ronquillo</surname><given-names>CE</given-names> </name><name name-style="western"><surname>Peltonen</surname><given-names>LM</given-names> </name><name name-style="western"><surname>Pruinelli</surname><given-names>L</given-names> </name><etal/></person-group><article-title>Artificial intelligence in nursing: priorities and opportunities from an international invitational think-tank of the Nursing and Artificial Intelligence Leadership Collaborative</article-title><source>J Adv Nurs</source><year>2021</year><month>09</month><volume>77</volume><issue>9</issue><fpage>3707</fpage><lpage>3717</lpage><pub-id pub-id-type="doi">10.1111/jan.14855</pub-id><pub-id pub-id-type="medline">34003504</pub-id></nlm-citation></ref><ref id="ref4"><label>4</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>von Gerich</surname><given-names>H</given-names> </name><name name-style="western"><surname>Moen</surname><given-names>H</given-names> </name><name name-style="western"><surname>Block</surname><given-names>LJ</given-names> </name><etal/></person-group><article-title>Artificial intelligence-based technologies in nursing: a scoping literature review of the evidence</article-title><source>Int J Nurs Stud</source><year>2022</year><month>03</month><volume>127</volume><fpage>104153</fpage><pub-id pub-id-type="doi">10.1016/j.ijnurstu.2021.104153</pub-id><pub-id pub-id-type="medline">35092870</pub-id></nlm-citation></ref><ref id="ref5"><label>5</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Cabitza</surname><given-names>F</given-names> </name><name name-style="western"><surname>Rasoini</surname><given-names>R</given-names> </name><name name-style="western"><surname>Gensini</surname><given-names>GF</given-names> </name></person-group><article-title>Unintended consequences of machine learning in medicine</article-title><source>JAMA</source><year>2017</year><month>08</month><day>8</day><volume>318</volume><issue>6</issue><fpage>517</fpage><lpage>518</lpage><pub-id pub-id-type="doi">10.1001/jama.2017.7797</pub-id><pub-id pub-id-type="medline">28727867</pub-id></nlm-citation></ref><ref id="ref6"><label>6</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Amann</surname><given-names>J</given-names> </name><name name-style="western"><surname>Blasimme</surname><given-names>A</given-names> </name><name name-style="western"><surname>Vayena</surname><given-names>E</given-names> </name><name name-style="western"><surname>Frey</surname><given-names>D</given-names> </name><name name-style="western"><surname>Madai</surname><given-names>VI</given-names> </name><collab>Precise4Q consortium</collab></person-group><article-title>Explainability for artificial intelligence in healthcare: a multidisciplinary perspective</article-title><source>BMC Med Inform Decis Mak</source><year>2020</year><month>11</month><day>30</day><volume>20</volume><issue>1</issue><fpage>310</fpage><pub-id pub-id-type="doi">10.1186/s12911-020-01332-6</pub-id><pub-id pub-id-type="medline">33256715</pub-id></nlm-citation></ref><ref id="ref7"><label>7</label><nlm-citation citation-type="confproc"><person-group person-group-type="author"><name name-style="western"><surname>Bussone</surname><given-names>A</given-names> </name><name name-style="western"><surname>Stumpf</surname><given-names>S</given-names> </name><name name-style="western"><surname>O&#x2019;Sullivan</surname><given-names>D</given-names> </name></person-group><article-title>The role of explanations on trust and reliance in clinical decision support systems</article-title><conf-name>2015 International Conference on Healthcare Informatics (ICHI)</conf-name><conf-date>Oct 21-23, 2015</conf-date><pub-id pub-id-type="doi">10.1109/ICHI.2015.26</pub-id></nlm-citation></ref><ref id="ref8"><label>8</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Sutton</surname><given-names>RT</given-names> </name><name name-style="western"><surname>Pincock</surname><given-names>D</given-names> </name><name name-style="western"><surname>Baumgart</surname><given-names>DC</given-names> </name><name name-style="western"><surname>Sadowski</surname><given-names>DC</given-names> </name><name name-style="western"><surname>Fedorak</surname><given-names>RN</given-names> </name><name name-style="western"><surname>Kroeker</surname><given-names>KI</given-names> </name></person-group><article-title>An overview of clinical decision support systems: benefits, risks, and strategies for success</article-title><source>NPJ Digit Med</source><year>2020</year><volume>3</volume><fpage>17</fpage><pub-id pub-id-type="doi">10.1038/s41746-020-0221-y</pub-id><pub-id pub-id-type="medline">32047862</pub-id></nlm-citation></ref><ref id="ref9"><label>9</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Parasuraman</surname><given-names>R</given-names> </name><name name-style="western"><surname>Riley</surname><given-names>V</given-names> </name></person-group><article-title>Humans and automation: use, misuse, disuse, abuse</article-title><source>Hum Factors</source><year>1997</year><month>06</month><volume>39</volume><issue>2</issue><fpage>230</fpage><lpage>253</lpage><pub-id pub-id-type="doi">10.1518/001872097778543886</pub-id></nlm-citation></ref><ref id="ref10"><label>10</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Lee</surname><given-names>JD</given-names> </name><name name-style="western"><surname>See</surname><given-names>KA</given-names> </name></person-group><article-title>Trust in automation: designing for appropriate reliance</article-title><source>Hum Factors</source><year>2004</year><volume>46</volume><issue>1</issue><fpage>50</fpage><lpage>80</lpage><pub-id pub-id-type="doi">10.1518/hfes.46.1.50_30392</pub-id><pub-id pub-id-type="medline">15151155</pub-id></nlm-citation></ref><ref id="ref11"><label>11</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Goddard</surname><given-names>K</given-names> </name><name name-style="western"><surname>Roudsari</surname><given-names>A</given-names> </name><name name-style="western"><surname>Wyatt</surname><given-names>JC</given-names> </name></person-group><article-title>Automation bias: a systematic review of frequency, effect mediators, and mitigators</article-title><source>J Am Med Inform Assoc</source><year>2012</year><volume>19</volume><issue>1</issue><fpage>121</fpage><lpage>127</lpage><pub-id pub-id-type="doi">10.1136/amiajnl-2011-000089</pub-id><pub-id pub-id-type="medline">21685142</pub-id></nlm-citation></ref><ref id="ref12"><label>12</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Lyell</surname><given-names>D</given-names> </name><name name-style="western"><surname>Coiera</surname><given-names>E</given-names> </name></person-group><article-title>Automation bias and verification complexity: a systematic review</article-title><source>J Am Med Inform Assoc</source><year>2017</year><month>03</month><day>1</day><volume>24</volume><issue>2</issue><fpage>423</fpage><lpage>431</lpage><pub-id pub-id-type="doi">10.1093/jamia/ocw105</pub-id><pub-id pub-id-type="medline">27516495</pub-id></nlm-citation></ref><ref id="ref13"><label>13</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Bu&#x00E7;inca</surname><given-names>Z</given-names> </name><name name-style="western"><surname>Malaya</surname><given-names>MB</given-names> </name><name name-style="western"><surname>Gajos</surname><given-names>KZ</given-names> </name></person-group><article-title>To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making</article-title><source>Proc ACM Hum-Comput Interact</source><year>2021</year><month>04</month><day>13</day><volume>5</volume><issue>CSCW1</issue><fpage>1</fpage><lpage>21</lpage><pub-id pub-id-type="doi">10.1145/3449287</pub-id></nlm-citation></ref><ref id="ref14"><label>14</label><nlm-citation citation-type="confproc"><person-group person-group-type="author"><name name-style="western"><surname>Yin</surname><given-names>M</given-names> </name><name name-style="western"><surname>Wortman Vaughan</surname><given-names>J</given-names> </name><name name-style="western"><surname>Wallach</surname><given-names>H</given-names> </name></person-group><article-title>Understanding the effect of accuracy on trust in machine learning models</article-title><conf-name>Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI &#x2019;19)</conf-name><conf-date>May 4-9, 2019</conf-date><pub-id pub-id-type="doi">10.1145/3290605.3300509</pub-id></nlm-citation></ref><ref id="ref15"><label>15</label><nlm-citation citation-type="confproc"><person-group person-group-type="author"><name name-style="western"><surname>Lu</surname><given-names>Z</given-names> </name><name name-style="western"><surname>Yin</surname><given-names>M</given-names> </name></person-group><article-title>Human reliance on machine learning models when performance feedback is limited: heuristics and risks</article-title><conf-name>Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI &#x2019;21)</conf-name><conf-date>May 8-13, 2021</conf-date><pub-id pub-id-type="doi">10.1145/3411764.3445562</pub-id></nlm-citation></ref><ref id="ref16"><label>16</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>K&#x00FC;cking</surname><given-names>F</given-names> </name><name name-style="western"><surname>Busch</surname><given-names>DA</given-names> </name><name name-style="western"><surname>Przysucha</surname><given-names>M</given-names> </name><etal/></person-group><article-title>Impact of AI recommendation correctness on diagnostic accuracy in clinical decision-making</article-title><source>Int J Med Inform</source><year>2026</year><month>03</month><day>1</day><volume>207</volume><fpage>106223</fpage><pub-id pub-id-type="doi">10.1016/j.ijmedinf.2025.106223</pub-id><pub-id pub-id-type="medline">41391283</pub-id></nlm-citation></ref><ref id="ref17"><label>17</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Singhal</surname><given-names>K</given-names> </name><name name-style="western"><surname>Azizi</surname><given-names>S</given-names> </name><name name-style="western"><surname>Tu</surname><given-names>T</given-names> </name><etal/></person-group><article-title>Large language models encode clinical knowledge</article-title><source>Nature</source><year>2023</year><month>08</month><volume>620</volume><issue>7972</issue><fpage>172</fpage><lpage>180</lpage><pub-id pub-id-type="doi">10.1038/s41586-023-06291-2</pub-id><pub-id pub-id-type="medline">37438534</pub-id></nlm-citation></ref><ref id="ref18"><label>18</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Hager</surname><given-names>P</given-names> </name><name name-style="western"><surname>Jungmann</surname><given-names>F</given-names> </name><name name-style="western"><surname>Holland</surname><given-names>R</given-names> </name><etal/></person-group><article-title>Evaluation and mitigation of the limitations of large language models in clinical decision-making</article-title><source>Nat Med</source><year>2024</year><month>09</month><volume>30</volume><issue>9</issue><fpage>2613</fpage><lpage>2622</lpage><pub-id pub-id-type="doi">10.1038/s41591-024-03097-1</pub-id><pub-id pub-id-type="medline">38965432</pub-id></nlm-citation></ref><ref id="ref19"><label>19</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Eriksen</surname><given-names>AV</given-names> </name><name name-style="western"><surname>M&#x00F6;ller</surname><given-names>S</given-names> </name><name name-style="western"><surname>Ryg</surname><given-names>J</given-names> </name></person-group><article-title>Use of GPT-4 to diagnose complex clinical cases</article-title><source>NEJM AI</source><year>2024</year><month>01</month><volume>1</volume><issue>1</issue><fpage>AIp2300031</fpage><pub-id pub-id-type="doi">10.1056/AIp2300031</pub-id></nlm-citation></ref><ref id="ref20"><label>20</label><nlm-citation citation-type="web"><person-group person-group-type="author"><name name-style="western"><surname>Lu</surname><given-names>Z</given-names> </name><name name-style="western"><surname>Yin</surname><given-names>M</given-names> </name></person-group><article-title>Trustworthy-ML: data and analysis code for &#x201C;Human Reliance on Machine Learning Models When Performance Feedback is Limited: Heuristics and Risks&#x201D;</article-title><source>GitHub</source><year>2021</year><access-date>2026-07-30</access-date><comment><ext-link ext-link-type="uri" xlink:href="https://github.com/ZhuoranLu/Trustworthy-ML">https://github.com/ZhuoranLu/Trustworthy-ML</ext-link></comment></nlm-citation></ref><ref id="ref21"><label>21</label><nlm-citation citation-type="book"><person-group person-group-type="author"><name name-style="western"><surname>Benner</surname><given-names>P</given-names> </name></person-group><source>From Novice to Expert: Excellence and Power in Clinical Nursing Practice</source><year>1984</year><publisher-name>Addison-Wesley</publisher-name><pub-id pub-id-type="other">9780201002997</pub-id></nlm-citation></ref><ref id="ref22"><label>22</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Benner</surname><given-names>P</given-names> </name></person-group><article-title>Using the Dreyfus model of skill acquisition to describe and interpret skill acquisition and clinical judgment in nursing practice and education</article-title><source>Bull Sci Technol Soc</source><year>2004</year><month>06</month><volume>24</volume><issue>3</issue><fpage>188</fpage><lpage>199</lpage><pub-id pub-id-type="doi">10.1177/0270467604265061</pub-id></nlm-citation></ref><ref id="ref23"><label>23</label><nlm-citation citation-type="confproc"><person-group person-group-type="author"><name name-style="western"><surname>Bansal</surname><given-names>G</given-names> </name><name name-style="western"><surname>Wu</surname><given-names>T</given-names> </name><name name-style="western"><surname>Zhou</surname><given-names>J</given-names> </name><etal/></person-group><article-title>Does the whole exceed its parts? the effect of AI explanations on complementary team performance</article-title><conf-name>Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI &#x2019;21)</conf-name><conf-date>May 8-13, 2021</conf-date><pub-id pub-id-type="doi">10.1145/3411764.3445717</pub-id></nlm-citation></ref><ref id="ref24"><label>24</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Vaccaro</surname><given-names>M</given-names> </name><name name-style="western"><surname>Almaatouq</surname><given-names>A</given-names> </name><name name-style="western"><surname>Malone</surname><given-names>T</given-names> </name></person-group><article-title>When combinations of humans and AI are useful: a systematic review and meta-analysis</article-title><source>Nat Hum Behav</source><year>2024</year><month>12</month><volume>8</volume><issue>12</issue><fpage>2293</fpage><lpage>2303</lpage><pub-id pub-id-type="doi">10.1038/s41562-024-02024-1</pub-id><pub-id pub-id-type="medline">39468277</pub-id></nlm-citation></ref><ref id="ref25"><label>25</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Goh</surname><given-names>E</given-names> </name><name name-style="western"><surname>Gallo</surname><given-names>R</given-names> </name><name name-style="western"><surname>Hom</surname><given-names>J</given-names> </name><etal/></person-group><article-title>Large language model influence on diagnostic reasoning: a randomized clinical trial</article-title><source>JAMA Netw Open</source><year>2024</year><month>10</month><day>1</day><volume>7</volume><issue>10</issue><fpage>e2440969</fpage><pub-id pub-id-type="doi">10.1001/jamanetworkopen.2024.40969</pub-id><pub-id pub-id-type="medline">39466245</pub-id></nlm-citation></ref><ref id="ref26"><label>26</label><nlm-citation citation-type="confproc"><person-group person-group-type="author"><name name-style="western"><surname>Schemmer</surname><given-names>M</given-names> </name><name name-style="western"><surname>Kuehl</surname><given-names>N</given-names> </name><name name-style="western"><surname>Benz</surname><given-names>C</given-names> </name><name name-style="western"><surname>Bartos</surname><given-names>A</given-names> </name><name name-style="western"><surname>Satzger</surname><given-names>G</given-names> </name></person-group><article-title>Appropriate reliance on AI advice: conceptualization and the effect of explanations</article-title><conf-name>Proceedings of the 28th International Conference on Intelligent User Interfaces (IUI &#x2019;23)</conf-name><conf-date>Mar 27-31, 2023</conf-date><pub-id pub-id-type="doi">10.1145/3581641.3584066</pub-id></nlm-citation></ref><ref id="ref27"><label>27</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>K&#x00FC;cking</surname><given-names>F</given-names> </name><name name-style="western"><surname>H&#x00FC;bner</surname><given-names>UH</given-names> </name><name name-style="western"><surname>Busch</surname><given-names>D</given-names> </name></person-group><article-title>Diagnostic accuracy differences in detecting wound maceration between humans and artificial intelligence: the role of human expertise revisited</article-title><source>J Am Med Inform Assoc</source><year>2025</year><month>09</month><day>1</day><volume>32</volume><issue>9</issue><fpage>1425</fpage><lpage>1433</lpage><pub-id pub-id-type="doi">10.1093/jamia/ocaf116</pub-id><pub-id pub-id-type="medline">40668943</pub-id></nlm-citation></ref><ref id="ref28"><label>28</label><nlm-citation citation-type="book"><person-group person-group-type="author"><name name-style="western"><surname>Kahneman</surname><given-names>D</given-names> </name></person-group><source>Thinking, Fast and Slow</source><year>2011</year><publisher-name>Farrar, Straus and Giroux</publisher-name><pub-id pub-id-type="other">9780374275631</pub-id></nlm-citation></ref><ref id="ref29"><label>29</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Buchanan</surname><given-names>C</given-names> </name><name name-style="western"><surname>Howitt</surname><given-names>ML</given-names> </name><name name-style="western"><surname>Wilson</surname><given-names>R</given-names> </name><name name-style="western"><surname>Booth</surname><given-names>RG</given-names> </name><name name-style="western"><surname>Risling</surname><given-names>T</given-names> </name><name name-style="western"><surname>Bamford</surname><given-names>M</given-names> </name></person-group><article-title>Predicted influences of artificial intelligence on nursing education: scoping review</article-title><source>JMIR Nurs</source><year>2021</year><volume>4</volume><issue>1</issue><fpage>e23933</fpage><pub-id pub-id-type="doi">10.2196/23933</pub-id><pub-id pub-id-type="medline">34345794</pub-id></nlm-citation></ref><ref id="ref30"><label>30</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Gilbert</surname><given-names>S</given-names> </name><name name-style="western"><surname>Harvey</surname><given-names>H</given-names> </name><name name-style="western"><surname>Melvin</surname><given-names>T</given-names> </name><name name-style="western"><surname>Vollebregt</surname><given-names>E</given-names> </name><name name-style="western"><surname>Wicks</surname><given-names>P</given-names> </name></person-group><article-title>Large language model AI chatbots require approval as medical devices</article-title><source>Nat Med</source><year>2023</year><month>10</month><volume>29</volume><issue>10</issue><fpage>2396</fpage><lpage>2398</lpage><pub-id pub-id-type="doi">10.1038/s41591-023-02412-6</pub-id><pub-id pub-id-type="medline">37391665</pub-id></nlm-citation></ref><ref id="ref31"><label>31</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Mesk&#x00F3;</surname><given-names>B</given-names> </name><name name-style="western"><surname>Topol</surname><given-names>EJ</given-names> </name></person-group><article-title>The imperative for regulatory oversight of large language models (or generative AI) in healthcare</article-title><source>NPJ Digit Med</source><year>2023</year><month>07</month><day>6</day><volume>6</volume><issue>1</issue><fpage>120</fpage><pub-id pub-id-type="doi">10.1038/s41746-023-00873-0</pub-id><pub-id pub-id-type="medline">37414860</pub-id></nlm-citation></ref><ref id="ref32"><label>32</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>K&#x00FC;per</surname><given-names>A</given-names> </name><name name-style="western"><surname>Lodde</surname><given-names>GC</given-names> </name><name name-style="western"><surname>Livingstone</surname><given-names>E</given-names> </name><name name-style="western"><surname>Schadendorf</surname><given-names>D</given-names> </name><name name-style="western"><surname>Kr&#x00E4;mer</surname><given-names>N</given-names> </name></person-group><article-title>Psychological factors influencing appropriate reliance on AI-enabled clinical decision support systems: experimental web-based study among dermatologists</article-title><source>J Med Internet Res</source><year>2025</year><month>04</month><day>4</day><volume>27</volume><fpage>e58660</fpage><pub-id pub-id-type="doi">10.2196/58660</pub-id><pub-id pub-id-type="medline">40184614</pub-id></nlm-citation></ref><ref id="ref33"><label>33</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>K&#x00FC;cking</surname><given-names>F</given-names> </name><name name-style="western"><surname>H&#x00FC;bner</surname><given-names>U</given-names> </name><name name-style="western"><surname>Przysucha</surname><given-names>M</given-names> </name><etal/></person-group><article-title>Automation bias in AI-decision support: results from an empirical study</article-title><source>Stud Health Technol Inform</source><year>2024</year><month>08</month><day>30</day><volume>317</volume><fpage>298</fpage><lpage>304</lpage><pub-id pub-id-type="doi">10.3233/SHTI240871</pub-id><pub-id pub-id-type="medline">39234734</pub-id></nlm-citation></ref><ref id="ref34"><label>34</label><nlm-citation citation-type="web"><person-group person-group-type="author"><name name-style="western"><surname>Tajima</surname><given-names>H</given-names> </name></person-group><article-title>Nursing-AI-error-thresholds: reproducibility package (calibration, bootstrap, 27-cell simulation, joint model, and sensitivity analyses)</article-title><source>GitHub</source><year>2026</year><access-date>2026-07-30</access-date><comment><ext-link ext-link-type="uri" xlink:href="https://github.com/aiprof202604-tech/nursing-ai-error-thresholds">https://github.com/aiprof202604-tech/nursing-ai-error-thresholds</ext-link></comment></nlm-citation></ref><ref id="ref35"><label>35</label><nlm-citation citation-type="other"><person-group person-group-type="author"><name name-style="western"><surname>Tajima</surname><given-names>H</given-names> </name></person-group><article-title>Theoretical exploration of error thresholds for clinical AI decision support in nursing: an exploratory simulation study grounded in human-AI reliance data</article-title><source>SSRN</source><comment>Preprint posted online on  Apr 24, 2026</comment><pub-id pub-id-type="doi">10.2139/ssrn.6632719</pub-id></nlm-citation></ref></ref-list><app-group><supplementary-material id="app1"><label>Multimedia Appendix 1</label><p>Extended robustness analyses. (A) Minimum AI accuracy (<italic>A</italic><sub>threshold</sub>) required to keep predicted team error below the 10% and 20% targets under the joint human-AI decision model, plotted against the clinician&#x2019;s independent accuracy <italic>p</italic><sub>c</sub>, for the highest-risk (novice&#x00D7;high-complexity) and lowest-deference (experienced&#x00D7;low-complexity) cells. The primary model corresponds to <italic>p</italic><sub>c</sub>&#x2192;1. In the lowest-deference cell, the team error acquires an irreducible floor that grows as <italic>p</italic><sub>c</sub> declines, so the 10% target becomes unreachable by any AI accuracy at <italic>p</italic><sub>c</sub>&#x2264;0.7. (B) Distribution of the novice&#x00D7;high-complexity <italic>A</italic><sub>threshold</sub> from the global Monte Carlo sensitivity analysis (50,000 draws per prior) for the &#x003C;10% and &#x003C;20% targets, under a theory-signed prior and a fully sign-agnostic prior. The vertical line marks the reported estimate (0.894) and the dotted line the 0.70 anchor.</p><media xlink:href="nursing_v9i1e99590_app1.png" xlink:title="PNG File, 209 KB"/></supplementary-material><supplementary-material id="app2"><label>Multimedia Appendix 2</label><p>Calibration of the linear reliance model against 9 empirical data points from 3 independent studies (N=3502).</p><media xlink:href="nursing_v9i1e99590_app2.png" xlink:title="PNG File, 176 KB"/></supplementary-material></app-group></back></article>