Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract Background Large language models (LLMs) are emerging as tools to support clinical decision making. HIV management is a compelling use case due to its complexity and dynamic nature, involving diverse treatment options, comorbidities, and adherence challenges. However, integrating LLMs into clinical practice raises concerns about accuracy, safety, and clinician acceptance. Despite growing interest, their performance in HIV care remains poorly studied, and benchmarking is lacking. Methods We developed HIVMedQA, a clinician-curated benchmark of HIV-related open-ended medical question-answer pairs spanning basic knowledge, clinical reasoning, complex patient vignettes, and bias-modified scenarios. We evaluated seven general-purpose and three medical LLMs. Performance was assessed using lexical similarity and an extended medical LLM-as-a-judge framework capturing key clinical dimensions, including question comprehension, reasoning, knowledge recall, bias, potential harm, and factual accuracy, to better capture nuances relevant to the medical domain, with additional evaluation by HIV-experienced physicians. Results Performance varies substantially across models and task complexity. Gemini 2.5 Pro achieves the highest overall scores, followed by Claude 3.5 Sonnet and MedGemma-27B. Knowledge recall is generally stronger than question comprehension or clinical reasoning. Medical LLMs do not consistently outperform general-purpose models, and model size alone does not predict performance. Several models are sensitive to cognitive bias prompts. LLM-as-a-judge scoring aligns better with clinician assessment than lexical metrics. Conclusions HIVMedQA provides a structured benchmark for evaluating LLMs in HIV clinical decision support. Current LLMs show promise, but limitations in reasoning, bias robustness, and safety indicate that careful validation, domain-specific evaluation, and clinician oversight remain essential before clinical deployment.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Gonzalo Cardenal-Antolin, Jacques Fellay, Bashkim Jaha, Huldrych F. Günthard, Johannes Nemeth, Roger Kouyos, Niko Beerenwinkel, Diane Duroux
- Quelle
- Communications Medicine
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2730-664X
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Gonzalo Cardenal-Antolin, Jacques Fellay, Bashkim Jaha, Huldrych F. Günthard, Johannes Nemeth, Roger Kouyos, Niko Beerenwinkel, Diane Duroux (2026). Benchmarking large language models for HIV medical decision support. Communications Medicine. https://doi.org/10.1038/s43856-026-01875-1
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1