Vollständiger Abstract
Worum geht es in dieser Arbeit?
Speech Emotion Recognition (SER) is an important component of human-centered intelligent systems, yet robust performance remains challenging when speaker identities differ between training and testing. This study presents a protocol-aware and reproducible comparison on the CREMA-D corpus using three pipelines: (i) a classical MFCC-based Support Vector Machine (SVM), (ii) a log-mel Convolutional Neural Network (CNN), and (iii) a lightweight hybrid model that concatenates handcrafted acoustic descriptors with CNN-derived embeddings and uses an SVM classifier. The methodological contribution is not a new standalone classifier; it is the controlled integration of identical preprocessing, random and actor-wise evaluation, five-seed robustness reporting, class-wise error analysis, and CPU-oriented deployment within one experimental framework. The Hybrid approach achieves the best overall performance, obtaining 62.03% ± 0.94% Macro-F1 on the random split and 58.09% ± 1.36% on the actor-wise split, outperforming MFCC+SVM (55.72% ± 0.95% and 52.05% ± 1.95%) and the log-mel CNN (50.20% ± 1.51% and 42.68% ± 1.84%). A Streamlit interface supports WAV upload, live prediction, and export of per-seed confusion matrices and summary figures. The results show that a controlled lightweight fusion framework can improve robustness while making the performance gap between speaker-overlapping and speaker-independent evaluation explicit.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Parveen Kumari, Yogita Yashveer Raghav, Vimmi Kochher, Prashanth kumar katta, Abhilasha A, Prabhakar M, Jayanthiladevi A, Fikir Gizachew Belete
- Quelle
- PLOS One
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 1932-6203
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Parveen Kumari, Yogita Yashveer Raghav, Vimmi Kochher, Prashanth kumar katta, Abhilasha A, Prabhakar M, Jayanthiladevi A, Fikir Gizachew Belete (2026). Hybrid CNN-embedding fusion with MFCC-SVM for speech emotion recognition: Random vs actor-wise evaluation on CREMA-D. PLOS One. https://doi.org/10.1371/journal.pone.0355238
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1