Vollständiger Abstract
Worum geht es in dieser Arbeit?
Speech emotion recognition (SER) infers affective states from speech signals, but positional encoding for acoustic tokens remains underexplored in Transformer-based SER. Existing models often reuse encodings designed for text and do not explicitly account for the different sequential and two-dimensional structures of acoustic representations. We propose Fibonacci Position Embedding (FPE) and Fibonacci Target Shutter (FTS). FTS constructs overlapping candidate-index sets over time-frequency token grids, and FPE samples a Fibonacci index and applies its modulo-wrapped, dimension-dependent phase rotation to query and key vectors. The modules are integrated into STRE-Former, which fuses Wav2Vec, log-mel spectrogram, and MFCC representations through asymmetric cross-representation attention with representation-specific positional encodings. We also introduce an implementation-consistent conditional-entropy formulation that quantifies uncertainty in recovering a token location from its sampled positional representation; this quantity characterizes positional ambiguity rather than downstream modeling capacity. Experiments over 64 positional-encoding combinations on IEMOCAP and MELD identify dataset-dependent highest-observed configurations, reaching 74.21% weighted accuracy on IEMOCAP-4, 74.54% on IEMOCAP-6, and 49.44% on MELD. These empirical observations suggest that the relative behavior of positional-encoding strategies may depend on the acoustic representation and evaluation dataset, rather than supporting a single universally optimal scheme.
Abstract: PubMed · Datensatz
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Nicht angegeben
- Quelle
- CrossRef Listing of Deleted DOIs
- Publikation
- 2000-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0849-6757
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
(2000). 10.3390/polym8030084. CrossRef Listing of Deleted DOIs. https://doi.org/10.3390/s26165050