Vollständiger Abstract
Worum geht es in dieser Arbeit?
Speech deepfake detection (SDD) is commonly formulated as offline utterance-level binary classification, which limits early decisions in streaming communication and provides little mechanism-level evidence for a spoofing prediction. We propose C-MCSS-Mamba, a block-causal SDD framework that processes fixed-duration audio blocks with a partially fine-tuned XLS-R frontend and carries detection states across blocks through Mamba. Its selective scan is replaced by the proposed Counterfactual Mechanism-Contrastive Selective Scan (C-MCSS), which maintains shared, text-to-speech (TTS), and voice conversion (VC) states. Frame-level TTS/VC hypotheses directly modulate state update, retention, and forgetting. A Causal Counterfactual Evidence Reliability Module (C-CERM) further suppresses transient mechanism evidence through a recurrent reliability state that adaptively regulates these dynamics. Model-internal TTS/VC evidence is obtained by counterfactually disabling the corresponding recurrent state and measuring the resulting spoof-logit decrease. The model produces prefix-level spoof decisions and contribution-based state-dependence measures; these measures only characterize the model’s internal decision process, and do not identify the physical speech generation process of an individual sample. A joint objective combines detection, early-prefix, mechanism, contrastive, and intervention supervision. Experiments demonstrate competitive detection performance, achieving 0.47% EER on the ASVspoof 2019 LA evaluation set while maintaining effectiveness across four cross-dataset benchmarks.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Gaopeng Zhang, Shidong Liu, Dengtao Zhang, Liang Tang
- Quelle
- Applied Sciences
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2076-3417
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Gaopeng Zhang, Shidong Liu, Dengtao Zhang, Liang Tang (2026). C-MCSS-Mamba: Counterfactual Mechanism-Contrastive Selective Scan Within Mamba for Block-Causal Speech Deepfake Detection. Applied Sciences. https://doi.org/10.3390/app16178413
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1