Vollständiger Abstract
Worum geht es in dieser Arbeit?
Often, the initial data sets for machine learning contain an excessive number of features, among which there may be highly informative and insignificant, correlating or even noise variables. Solving the feature selection problem not only optimizes machine learning processes, but also opens up new opportunities for implementing machine learning in practice-oriented areas that require high accuracy and trust in algorithms. This phenomenon leads to a number of critical problems: distorted training of the model based on irrelevant patterns, a decrease in its generalizing ability, a sharp increase in computational costs for training and, as a result, difficulty in interpreting the results obtained. The paper proposes a method for optimizing the feature space which uses a combined approach that includes maximizing the measure of model effectiveness, selecting the highest-quality features based on information content and correlation analysis. In addition to optimizing the feature space, the model offers the best classifier for the data set used. As a solution to the objective function, a genetic algorithm with elitism is used to find the optimal value. The proposed model is demonstrated on a dataset in the field of medical diagnostics in comparison with the well-known method of recursive feature selection (RFE), correlation analysis, as well as with the results obtained using the full dataset. The data set consists of measurements of mammary glands using the method of microwave radiothermometry and labels characterizing the severity of temperature anomalies. The dataset contains 62 features and 6 labels and 9,310 measurement records. The results demonstrate the high efficiency of the proposed model either close to or exceeding the value of RFE. As a result of optimizing the feature space for the data set, it was possible to reduce from 62 features to 15, while not only not reducing the accuracy of the model, but even slightly increasing it, the best model of the classifier turned out to be the logistic regression model. Thus, the accuracy indicators were 0.79 for the complete data set, 0.7879 for 15 optimized features based on the RFE method, 0.7175 for 29 features obtained as a result of correlation analysis, and 0.7911 for 15 features obtained using the proposed model. The proposed method not only effectively reduces the feature space, but also increases the accuracy of the classification model, despite a significant decrease in the number of features.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- K. S. Dyomin, I. V. Germashev
- Quelle
- Scientific and Technical Journal of Information Technologies, Mechanics and Optics
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2500-0373, 2226-1494
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
K. S. Dyomin, I. V. Germashev (2026). Combined feature selection in machine learning for disease diagnosis in medicine. Scientific and Technical Journal of Information Technologies, Mechanics and Optics. https://doi.org/10.17586/2226-1494-2026-26-4-763-770
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1