Vollständiger Abstract
Worum geht es in dieser Arbeit?
Introduction . The recognition of basic emotions from facial images is increasingly required in driver monitoring systems, medical user interfaces, and educational learning analytics. As of July 2024, such systems have been mandated by EU regulations. Among the methodological approaches in this area, neural-network architectures reach high accuracy but require large training samples and remain opaque, which is critical in safety applications. Classical descriptors LBP and HOG, in contrast, are computationally efficient and interpretable. However, their performance under comparable conditions remains underexplored: there are still questions regarding the stability of their metrics under cross-validation on compact datasets, as well as the nature of confusions between visually similar emotions. In this context, it was assumed that the near-perfect accuracy of HOG observed in such protocols reflected leakage of subject identities between the training and test sets, thereby overestimating the generalization ability of the descriptor. The objective of the study was to evaluate LBP and HOG on the JAFFE dataset in a single, reproducible pipeline, and determine the limits of their applicability. Materials and Methods . To address the stated tasks, the study was organized as a comparative empirical evaluation of two classical descriptors within a single reproducible protocol. The open JAFFE dataset (Lyons et al., 1998; 213 grayscale images of ten subjects with seven basic emotions) served as the experimental basis, with quality assessed under stratified five-fold cross-validation. The images underwent intensity normalization (histogram equalization) and geometric alignment by facial landmarks. Then, on the prepared frames, the LBP (texture) and HOG (contour geometry) descriptors were extracted, with specific parameters reported in the body of the article. The resulting features were fed to a linear SVM (support vector machine), with the regularization parameter tuned using a grid search over {0.01, 0.1, 1, 10} via internal three-fold cross-validation. Quality was assessed as the mean ± standard deviation of Accuracy and Macro-F1 metrics across five outer folds. The software implementation of the entire pipeline was done in Python 3.10 (using the scikit-learn 1.3 and scikit-image 0.21 libraries). Results . In the experiment, HOG combined with linear SVM provided complete separation of the seven emotional classes on the JAFFE dataset within the selected stratified protocol. Under the same algorithmic order, LBP showed significantly lower accuracy and revealed a specific structure of systematic confusion errors: the pairs “fear — surprise” and “sadness — neutral” provided consistently indistinguishable, while the category-wise distribution of metrics quantified the fundamental differences between contour and texture feature representations. Discussion . The results obtained indicate the decisive impact of the data splitting protocol on the final assessment of classical descriptors. Specifically, the perfect separability of HOG features is explained by preservation of subject identities across training and test sets, and is fully consistent with the known sensitivity of the gradient profile to individual facial features. With respect to LBP, the observed values fall within the range known from the reference work by Shan, Gong, and McOwan. However, the pattern of false positives documented in the present study is the first to be interpreted in detail through FACS description of overlapping sets of active facial muscles, which has previously been absent from published comparative reviews. Conclusion . The work has successfully solved the problems of comparative evaluation of LBP and HOG methods, construction of a fully reproducible software pipeline, and meaningful interpretation of automatic classification errors. Based on the data obtained, the HOG model combined with a linear SVM is suitable for laboratory tasks with a fixed set of subjects, whereas LBP in the same combination is justified as a component of hybrid architectures with convolutional or recurrent networks. In applied terms, the results can be used in the design of emotion recognition systems for medical interfaces and driver-monitoring systems. At the same time, the limitations of the work include the compact JAFFE dataset and the laboratory imaging conditions. Addressing these limitations, future work will involve moving to a more demanding protocol with full subject-out exclusion from the training set, as well as a comparative evaluation against current deep learning models.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- S. A. Filichkin, S. V. Vologdin
- Quelle
- Advanced Engineering Research (Rostov-on-Don)
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2687-1653
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
S. A. Filichkin, S. V. Vologdin (2026). A Study of Human Emotion Recognition Methods Based on Digital Image Analysis. Advanced Engineering Research (Rostov-on-Don). https://doi.org/10.23947/2687-1653-2026-26-3-2437
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1