Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Benchmarking large language models for HIV medical decision support

Gonzalo Cardenal-Antolin, Jacques Fellay, Bashkim Jaha, Huldrych F. Günthard, Johannes Nemeth, Roger Kouyos, Niko Beerenwinkel, Diane Duroux

Communications Medicine · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Abstract Background Large language models (LLMs) are emerging as tools to support clinical decision making. HIV management is a compelling use case due to its complexity and dynamic nature, involving diverse treatment options, comorbidities, and adherence challenges. However, integrating LLMs into clinical practice raises concerns about accuracy, safety, and clinician acceptance. Despite growing interest, their performance in HIV care remains poorly studied, and benchmarking is lacking. Methods We developed HIVMedQA, a clinician-curated benchmark of HIV-related open-ended medical question-answer pairs spanning basic knowledge, clinical reasoning, complex patient vignettes, and bias-modified scenarios. We evaluated seven general-purpose and three medical LLMs. Performance was assessed using lexical similarity and an extended medical LLM-as-a-judge framework capturing key clinical dimensions, including question comprehension, reasoning, knowledge recall, bias, potential harm, and factual accuracy, to better capture nuances relevant to the medical domain, with additional evaluation by HIV-experienced physicians. Results Performance varies substantially across models and task complexity. Gemini 2.5 Pro achieves the highest overall scores, followed by Claude 3.5 Sonnet and MedGemma-27B. Knowledge recall is generally stronger than question comprehension or clinical reasoning. Medical LLMs do not consistently outperform general-purpose models, and model size alone does not predict performance. Several models are sensitive to cognitive bias prompts. LLM-as-a-judge scoring aligns better with clinician assessment than lexical metrics. Conclusions HIVMedQA provides a structured benchmark for evaluating LLMs in HIV clinical decision support. Current LLMs show promise, but limitations in reasoning, bias robustness, and safety indicate that careful validation, domain-specific evaluation, and clinician oversight remain essential before clinical deployment.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Gonzalo Cardenal-Antolin, Jacques Fellay, Bashkim Jaha, Huldrych F. Günthard, Johannes Nemeth, Roger Kouyos, Niko Beerenwinkel, Diane Duroux
Quelle
Communications Medicine
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2730-664X
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Gonzalo Cardenal-Antolin, Jacques Fellay, Bashkim Jaha, Huldrych F. Günthard, Johannes Nemeth, Roger Kouyos, Niko Beerenwinkel, Diane Duroux (2026). Benchmarking large language models for HIV medical decision support. Communications Medicine. https://doi.org/10.1038/s43856-026-01875-1
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1