Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract Background Accurate pre‑operative diagnosis of vocal cord lesions remains challenging. We developed a multimodal deep learning model that integrates laryngoscopic images, voice recordings, and biochemical markers to improve diagnostic accuracy. Methods In this retrospective study (2020–2025), a total of 425 patients were enrolled. Of these, 374 patients treated between January 2020 and May 2025 formed the development cohort (contributing 1,947 aligned multimodal samples), and the remaining 51 consecutive cases treated between June 2025 and December 2025 were reserved as an independent temporal test set. Using an early‑fusion strategy and a Transformer encoder, we built a unified diagnostic framework. Model performance was evaluated by accuracy, average precision, AUC, and loss. Results During training, the model achieved a sample‑level training accuracy of 98.59% and a validation accuracy of 93.46% (monitored on 198 samples for model selection). For final reporting, all validation and test metrics were computed at the patient level after softmax averaging per patient: the validation set ( n = 37 patients) yielded an accuracy of 94.59% (35/37) and an AUC of 0.979; the independent temporal test set ( n = 51 patients) achieved an accuracy of 96.08% (49/51) and an AUC of 0.972, with only two misclassifications. Inference‑masking experiments showed that each modality contributed non‑redundant information; voice masking caused the largest drop (−30.44% accuracy), despite its low static importance weight. Conclusions Multimodal deep learning may enhance diagnostic accuracy for vocal cord lesions and could help reduce missed diagnoses. Our framework offers a potentially scalable approach for integrating heterogeneous clinical data; however, these findings are preliminary and require confirmation in larger multi‑center cohorts.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Xue Zhao, Shuang Li, Linlin Zheng, Tuanjie Wang, Haixian Guo, Dan Yu
- Quelle
- European Journal of Medical Research
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2047-783X
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Xue Zhao, Shuang Li, Linlin Zheng, Tuanjie Wang, Haixian Guo, Dan Yu (2026). Intelligent diagnosis of vocal cord lesions via multimodal deep learning: integrating laryngoscopy, voice, and biomarkers. European Journal of Medical Research. https://doi.org/10.1186/s40001-026-05099-w
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1