Vollständiger Abstract
Worum geht es in dieser Arbeit?
The growing volume of unstructured clinical and genetic text calls for automated processing methods that can meaningfully support clinical decision-making. Most existing work, however, treats the analysis of patient clinical descriptions and the interpretation of genetic reports as unrelated tasks, disregarding the sequence in which real diagnostic decisions unfold. We study two stage-specific families of classification tasks corresponding to the two points at which such decisions are taken—pre-genetic triage from clinical narratives, and post-genetic interpretation of completed genetic reports—and evaluate each family under a single leakage-controlled protocol using TF-IDF text representation with classical machine learning. All experiments use a corpus of 546 records with confirmed provenance, partitioned by a group-constrained master split (349 train/88 validation/109 test, seed 42) in which all 181 text-similarity candidate pairs are confined to a single partition. Task subsets contain 299 and 326 documents at the first stage and 375 and 461 at the second. Text representation relied on Word TF-IDF, Character TF-IDF and combined Word + Character TF-IDF with Logistic Regression and Linear SVM; every configuration was selected on validation Macro F1 alone, with the test partition opened only after selection. RuBERT-tiny2, XLM-RoBERTa base and PubMedBERT were fine-tuned as transformer baselines over three seeds each. Under this protocol, the first stage reached Macro F1 = 0.6876 (Accuracy = 0.8308, ROC-AUC = 0.8500) and the second-stage Macro F1 = 0.8225 (Accuracy = 0.8571, ROC-AUC = 0.9377) for the binary classification of diagnostic status. A controlled comparison holding the documents, split, vectorizer, classifier and hyperparameters fixed and varying only the label version increased Accuracy from 0.5738 to 0.8361 (exact McNemar p = 0.000145), whereas Macro F1 increased numerically from 0.5165 to 0.5966 but not significantly (paired bootstrap p = 0.354; 95% CI of the difference [−0.091, 0.229]). Two Russian and multilingual encoders collapsed to the majority class (Macro F1 = 0.4583), while an English biomedical encoder scored higher than the classical comparator (0.7705 ± 0.0119 against 0.6876) without reaching significance (paired bootstrap p = 0.463). Thus, the controlled experiment supports improved overall correctness, mainly associated with the majority class, but does not establish improved balanced class-wise performance at this sample size. In the second-stage leakage analysis, removing the diagnostic conclusion reduced Macro F1 from 0.8225 to 0.6478 (paired bootstrap p = 0.001), masking the exact label-generating rules reduced it to 0.7455 (p = 0.021), and combined masking reduced it to 0.6389 (p < 0.001). The unmasked Stage 2 score therefore depends materially on explicit report cues and should not be interpreted as evidence of independent diagnostic reasoning. The contribution of this work is a stage-oriented formulation of clinical genetic text classification in which the target variable of each task is aligned with the information available at the corresponding point of the diagnostic pathway, evaluated under a reproducible leakage-controlled protocol with bootstrap confidence intervals and paired significance testing. A unified single-model baseline on the joint target reached Macro F1 = 0.5238, below either stage-specific model on its own target. Because the labels are rule-derived rather than independently expert-validated, the corpus is small, and no external validation was performed, these results characterize what is achievable on this corpus rather than demonstrating clinical readiness.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Assem Shayakhmetova, Madina Sambetbayeva, Vladimir Barakhnin, Anar Sultangaziyeva, Nurzhan Mukazhanov, Ardak Batyrkhanov, Sandugash Serikbayeva, Raushan Begim
- Quelle
- Information
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2078-2489
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Assem Shayakhmetova, Madina Sambetbayeva, Vladimir Barakhnin, Anar Sultangaziyeva, Nurzhan Mukazhanov, Ardak Batyrkhanov, Sandugash Serikbayeva, Raushan Begim (2026). Stage-Oriented Text Classification for Russian-Language Clinical and Genetic Documents: Pre-Genetic Triage and Post-Genetic Report Interpretation. Information. https://doi.org/10.3390/info17090825
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1