Vollständiger Abstract
Worum geht es in dieser Arbeit?
Detecting audio-visual DeepFake (AVDeepFake) is becoming increasingly important as synthetic media tools become widely accessible and spread across consumer devices. In this study, we present an Detecting audio-visual DeepFakes (AV-DeepFakes) has become increasingly critical with the rapid proliferation of accessible synthetic media generation tools across consumer platforms. In this work, we propose a highperformance, deployment-efficient AV-DeepFake detection framework tailored for real-world consumer devices. The proposed model integrates a 3D convolutional visual encoder with a 2D convolutional audio encoder to learn synchronized multimodal representations, effectively capturing spatial, spectral, and prosodic inconsistencies inherent in manipulated content. To detect temporal forgeries, we introduce a bidirectional complementary boundary module that precisely localizes manipulation onsets and offsets. A cross-modal attention fusion mechanism aggregates modality-specific cues, while an uncertainty-aware gating strategy suppresses unreliable signals to improve robustness. Furthermore, a cross-modal discrepancy minimization loss encourages alignment for genuine samples while maximizing divergence for forged content, strengthening multimodal consistency learning. Extensive evaluations on FaceForensics++ and LAV-DF demonstrate the effectiveness of the proposed approach, achieving 97.1% AUC for clip-level detection and an 81.2% F1-score for temporal boundary localization, while reducing inference time by 5× compared with transformer-based methods.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Nasir Saleem, Adeel Hussain, Sami Bourouis, Sami Dhahbi, Afef Dhahbi, Ahmad Ali, Zhuoqi Zeng
- Quelle
- International Journal of Interactive Multimedia and Artificial Intelligence
- Publikation
- 2026-08-28
- Band / Ausgabe
- 10 / 1
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 1989-1660
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Nasir Saleem, Adeel Hussain, Sami Bourouis, Sami Dhahbi, Afef Dhahbi, Ahmad Ali, Zhuoqi Zeng (2026). AV-DeepFake-Net: Attention-Guided and Uncertainty-Aware Network for Audiovisual DeepFake Detection. International Journal of Interactive Multimedia and Artificial Intelligence, 10 (1). https://doi.org/10.9781/0m4b5v36
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1