Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Crossref · journal-article

Intelligent Audio-based Emotion Recognition in Speech by Deep Learning and Feature Engineering Techniques

Ramakrishna Gandi, A. Geetha, B. Ramasubba Reddy

International Journal of Computer Information Systems and Industrial Management Applications · 2026 · Band 18 · Ausgabe 19s · S. 14-33

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Speech Emotion Recognition (SER) is now progressively vital for many practical uses including virtual assistants, customer service, and healthcare monitoring as well as for SER systems still suffer with environmental noise, speaker variability, and cross-lingual adaptation that affect their accuracy and generalizing power even if they have made tremendous progress. This work introduces ExpressNet, an optimum Multi-Layer Perceptron (MLP)-based SER model aimed to solve these issues by leveraging a wide range of prosodic and spectral qualities incorporating Mel-Frequency Cepstral Coefficients (MFCCs), spectral contrast, and pitch variations. ReLU activation and a softmax output layer allow the model to classify six emotional states: anger, disgust, fear, happiness, neutral, and sad using a deep learning architecture. We assess ExpressNet using the CREMA-D dataset and achieve a test accuracy of 92.97%, above the results of previous state-of-the-art approaches. Particularly real-time applications gain from the method since it helps to blend high classification accuracy with computing efficiency. Our work emphasizes the need of applying deep learning methods with enhanced feature engineering to improve SER performance. Furthermore, we show a thorough assessment over numerous benchmark datasets to show the power and applicability capacity of the model in many different settings. Apart from being better than other options, ExpressNet is a consistent choice for use in real-world settings since it has a low overfitting rate. This work advances emotional computing by providing a solid and scalable foundation for SER, which will enable further research in emotional recognition systems. The probable utilization of this technique extends to mental health monitoring and human-computer interaction because it demonstrates excellence at handling complex emotional patterns in voice signals. Extensive research into self-supervised learning and multimodal data integration and cross-lingual adaptation will improve the model's potential across multiple application scenarios.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Ramakrishna Gandi, A. Geetha, B. Ramasubba Reddy
Quelle
International Journal of Computer Information Systems and Industrial Management Applications
Publikation
2026-08-23
Band / Ausgabe
18 / 19s
Seiten
14-33
ISSN / ISBN
2150-7988, 2150-7988
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Ramakrishna Gandi, A. Geetha, B. Ramasubba Reddy (2026). Intelligent Audio-based Emotion Recognition in Speech by Deep Learning and Feature Engineering Techniques. International Journal of Computer Information Systems and Industrial Management Applications, 18 (19s), 14-33. https://doi.org/10.70917/ijcisim-2026-4998
RIS BibTeX CSL-JSON