Vollständiger Abstract
Worum geht es in dieser Arbeit?
This paper presents a speech emotion recognition approach and its implementation on a field-programmable gate array (FPGA) with efficient neural network inference. The proposed approach employs a hybrid neural network architecture that integrates convolutional neural networks (CNNs), long short-term memory (LSTM) networks, and fully connected neural networks (FCNNs) to effectively capture both local and temporal features from speech signals. This architecture is well suited for modeling time-series data and is applied to speech signals in this paper. For FPGA implementation, one-dimensional CNNs (1D CNNs) are employed to extract features from Mel-frequency cepstral coefficients (MFCCs), thereby reducing computational complexity and improving the efficiency of CNN-based processing. The experiments are conducted using 10-fold cross-validation, achieving average recognition rates of 96.12%, 97.79%, 96.91%, 94.01%, and 98.09% on the RAVDESS, EMO-DB, IEMOCAP, BAUM-1s, and eNTERFACE’05 databases, respectively. According to the recognition accuracy, the proposed method performs better than existing related studies. The emotion recognition model is then deployed on an FPGA using the proposed inference algorithms, and the accuracies obtained on the FPGA are identical to those achieved on the PC. These results verify the successful deployment of the designed model on the FPGA platform.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Shing-Tai Pan, Han-Jui Wu
- Quelle
- Electronics
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2079-9292
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Shing-Tai Pan, Han-Jui Wu (2026). FPGA Chip-Based Mobile Sensor Design for Speech Emotions Recognition. Electronics. https://doi.org/10.3390/electronics15173832
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1