Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Toward safer digital sexual health communication: evaluating the public health reliability of large language model responses on sexually transmitted infections

Shucheng Zhang, Shuo Wang, Xiaoyue Sun, Zhuqing Li, Yang Lin, Jiale Geng, Xiaoqing Si

Frontiers in Public Health · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Background The global burden of sexually transmitted infections (STIs) continues to rise, yet stigma drives many to seek sensitive health information from artificial intelligence chatbots. The quality, safety, readability, and destigmatization of large language model (LLM) responses on sexual health remain under-evaluated, particularly for non-Western models. Methods This cross-sectional study constructed 30 standardized English queries (5 themes × 6 questions) informed by Google Trends, Centers for Disease Control and Prevention (CDC) guidelines, and patient education platforms, and submitted them to three LLMs—GPT-5.4, DeepSeek-V4-Pro, and Kimi K2.6—via official APIs, in duplicate (180 responses). Two dermatovenereologists (>10 years' experience) rated the responses while blinded to platform identity, across six dimensions (accuracy, completeness, safety, understandability, actionability, and destigmatization) using a five-point scale, supplemented by four readability metrics (FKGL, FRE, GFI, and SMOG). Primary analysis used linear mixed-effects models retaining all individual ratings; the original aggregated non-parametric pipeline was retained as a sensitivity analysis. Reliability used ICC and quadratic-weighted Cohen's κ with 95% CIs; correlations used query-level cluster bootstrap with false-discovery-rate control. Results Inter-platform differences were significant for accuracy (χ 2 (2) = 19.12, Holm-adjusted P < 0.001), understandability (χ 2 (2) = 13.54, P = 0.006), and destigmatization (χ 2 (2) = 9.16, P = 0.041). GPT-5.4 and Kimi K2.6 achieved the highest accuracy (both median 5.00), both significantly above DeepSeek-V4-Pro and not significantly different from each other. DeepSeek-V4-Pro and Kimi K2.6 both used significantly more destigmatizing language than GPT-5.4, with no significant difference between them. GPT-5.4 produced the least readable text on all four metrics (all adjusted P < 0.01). Inter-rater reliability was good to excellent (ICC(2, 1) = 0.79–0.95; quadratic-weighted κ = 0.79–0.95). Conclusions The three LLMs showed distinct profiles: GPT-5.4 and Kimi K2.6 achieved the highest accuracy, whereas GPT-5.4 produced the least readable outputs; DeepSeek-V4-Pro and Kimi K2.6 used more destigmatizing language. To our knowledge, this is the first study to quantify destigmatizing language as an explicit evaluation dimension for LLM-generated STI content, albeit as an exploratory measure pending formal content validation. Findings are specific to these models and queries and can inform quality standards and monitoring for safer AI-powered sexual health communication.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Shucheng Zhang, Shuo Wang, Xiaoyue Sun, Zhuqing Li, Yang Lin, Jiale Geng, Xiaoqing Si
Quelle
Frontiers in Public Health
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2296-2565
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Shucheng Zhang, Shuo Wang, Xiaoyue Sun, Zhuqing Li, Yang Lin, Jiale Geng, Xiaoqing Si (2026). Toward safer digital sexual health communication: evaluating the public health reliability of large language model responses on sexually transmitted infections. Frontiers in Public Health. https://doi.org/10.3389/fpubh.2026.1940111
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1