Vollständiger Abstract
Worum geht es in dieser Arbeit?
<h4>Background</h4>Immunotherapy plays an important role in bladder cancer care, requiring ongoing patient education, symptom monitoring, and communication of pathology- and biomarker-related information. Locally deployed large language models (LLMs) may support these nurse-led activities, but their safety and clinical usability remain uncertain.<h4>Method</h4>We conducted a human-in-the-loop evaluation of a locally deployed Qwen2-7B-Instruct model using structured Mandarin-language benchmarks. Five task categories were assessed: general immunotherapy education, pathology-informed communication, biomarker and precision pathology communication, symptom triage, and pathology-risk-informed triage. Outputs were reviewed by specialist nurses, with physician or pathologist adjudication for high-risk, discordant, or pathology-related cases. Evaluation dimensions included accuracy, completeness, readability, actionability, safety, pathology-feature recognition, uncertainty communication, red-flag recognition, escalation appropriateness, and revision burden.<h4>Results</h4>The locally deployed model Qwen2-7B achieved an overall score of 3.83 ± 0.23 for general patient education, performing best in disease understanding and self-management but less well in warning-sign education. In pathology-informed communication, the model achieved an overall score of 3.57 ± 0.22, with stronger performance in patient-friendly translation and structured pathology-feature recognition than in risk communication and follow-up guidance. Biomarker communication represented a more challenging task (overall score 3.26 ± 0.25), with frequent needs for uncertainty statements, clinician referral, and correction of biomarker overinterpretation. In symptom triage, agreement with expert triage levels was 75.9%, red-flag recognition was 71.1%, and escalation sensitivity was 77.1%; under-triage and unsafe reassurance remained observable, particularly in high-risk irAE scenarios. Incorporating pathology-derived risk context improved triage agreement from 52.5 to 70.0% and reduced under-triage, although safety-critical errors persisted. Across all tasks, only 3.7% of outputs required no revision, whereas 31.4% required major revision or withholding. The highest review burden occurred in biomarker communication and high-risk irAE scenarios.<h4>Conclusions</h4>A locally deployed Qwen2-7B-based LLM can support nurse-reviewed communication in bladder cancer immunotherapy, particularly for standardized education and explanation of structured pathology information. However, biomarker communication, pathology-derived risk interpretation, and immune-related symptom escalation remain safety-sensitive tasks requiring structured nurse oversight and clinician backup. These findings support locally deployed LLMs as human-in-the-loop nursing support tools rather than autonomous communication or triage systems within precision uro-oncology care.
Abstract: PubMed · Datensatz
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Nicht angegeben
- Quelle
- CrossRef Listing of Deleted DOIs
- Publikation
- 2000-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0849-6757
- Zitationen
- 14 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
(2000). 10.3389/fpsyg.2012.00132. CrossRef Listing of Deleted DOIs. https://doi.org/10.3389/fmed.2026.1903690