Vollständiger Abstract
Worum geht es in dieser Arbeit?
Background Internet-based cognitive behavioral therapy (iCBT) is an effective and scalable alternative to face-to-face psychotherapy, but its reach is constrained by the time therapists spend reviewing patient input and manually drafting written responses. Studies suggest that large language models (LLMs) may be capable of generating high-quality therapeutic text, with the potential to support therapists in delivering treatment. Their suitability as therapist-support tools in structured iCBT, however, remains insufficiently studied. Objective This study aims to assess the quality of LLM-generated iCBT responses to patient messages by comparing them to the quality of responses produced by humans. Methods In a preregistered blinded clinician rating experiment, experienced clinicians assessed the quality of human-produced vs LLM-generated therapist responses within a simulated iCBT treatment for functional somatic disorder. Raters were exposed to a stimulus material consisting of 5 fictitious patient messages, each paired with 1 human and 1 LLM-generated response. Raters assessed message/response pairs on 5 quality dimensions (overall quality, helpfulness, empathy, professionalism, and protocol adherence) and were asked to indicate the source of the response (human/LLM). Analyses were primarily descriptive, supplemented by exploratory statistical tests and descriptive thematic content analysis of open-ended text fields. The full preregistered study protocol is available at Open Science Framework. Results A total of 61 raters provided data, while 54 were eligible and included for analysis. Human- and LLM-generated responses were rated similarly across quality dimensions on a 1-5 scale: overall quality (LLM: mean 4.00, SD 0.54 vs human: mean 3.96, SD 0.53; d=0.06), helpfulness (LLM: mean 3.85, SD 0.57 vs human: mean 3.93, SD 0.49; d=0.13), professionalism (LLM: mean 4.25, SD 0.53 vs human: mean 4.11, SD 0.53; d=0.24), protocol adherence (LLM: mean 4.13, SD 0.52 vs human: mean 4.13, SD 0.54; d=0.03) and empathy (LLM: mean 4.31, SD 0.47 vs human: mean 4.08, SD 0.50; d=0.42). Raters correctly identified the source of human-generated responses (mean 79%, SD 19.65%) more accurately than LLM-generated responses (mean 63%, SD 21.30%). In all, 30/54 (55%) raters responded to one or more open text fields. Qualitative analysis indicated that LLM-generated responses were perceived as polished but also generic and at times excessively empathetic. Conclusions LLM-generated responses were judged to be of comparable quality to those written by human therapists, though qualitative feedback indicated they were at times generic and insufficiently challenging. These findings provide initial support for the feasibility of using LLMs as therapist-support tools in iCBT, but further research is needed to determine whether their integration yields tangible clinical and organizational benefits.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Thomas Tandrup Lamm, Arthur Bran Herbener, Oliver Rønn Christensen, Malene Flensborg Damholdt, Kaare Bro Wellnitz, Heidi Frølund Pedersen, Lisbeth Frostholm
- Quelle
- JMIR Mental Health
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2368-7959
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Thomas Tandrup Lamm, Arthur Bran Herbener, Oliver Rønn Christensen, Malene Flensborg Damholdt, Kaare Bro Wellnitz, Heidi Frølund Pedersen, Lisbeth Frostholm (2026). Large Language Models as a New Tool for Therapists in Internet-Based Cognitive Behavioral Therapy: Blinded Clinician Rating Pilot Experiment. JMIR Mental Health. https://doi.org/10.2196/96835