Frag' FlorenceEvidenz. Klar. Anwendbar.
Uhr 7/8Sources Journal Tree
Easy Demo

Crossref · journal-article

<b>Deep learning-based multi-speaker separation and speech enhancement for forensic audio analysis</b>

Oluwaseun Adeniyi Ojerinde, Rasheed Taiwo Mudasiru, Ramatu Abubakar, Benjamin Ilunuamie Alenoghena

Nature Journal of Emerging Sciences Technologies and Innovations · 2026 · Band 10 · Ausgabe 5 · S. 533-548

Vollständiger Abstract

Worum geht es in dieser Arbeit?

The increasing use of audio recordings in criminal investigations has created a growing demand for intelligent forensic audio analysis systems capable of recovering intelligible speech from acoustically challenging environments. Forensic recordings frequently contain overlapping speakers, background conversations, and environmental noise, making reliable speaker identification and evidence extraction difficult. This study proposed a hybrid deep learning framework for multi-speaker separation and speech enhancement by integrating Gated Convolutional Neural Networks (GCNNs) for speech source separation with Long Short-Term Memory (LSTM) networks for temporal speech enhancement. Unlike conventional approaches that treat speech separation and enhancement as independent tasks, the proposed framework jointly optimizes both processes within a unified architecture while preserving forensic audio integrity. The framework further integrates time-frequency masking, adaptive Wiener filtering, and spectral gain enhancement to suppress background noise while preserving speech fidelity and enhancing low-level background speech that may contain valuable forensic information. A custom dataset comprising 500 two-speaker conversations was developed and evaluated across diverse acoustic environments with signal-to-noise ratios ranging from −20 dB to +20 dB. Experimental results demonstrated an average Signal-to-Distortion Ratio (SDR) improvement of 9.4 dB, an average Signal-to-Noise Ratio (SNR) improvement of 5.2 dB, and a Diarization Error Rate (DER) of 6.4%. The framework also achieved an average processing latency of 1.2 seconds for a 20-second audio segment, indicating its suitability for near real-time forensic applications. The proposed framework substantially improves signal quality, speech intelligibility and speaker discrimination while preserving evidential integrity, thereby providing an effective decision-support tool for forensic audio analysis in criminal investigations.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Oluwaseun Adeniyi Ojerinde, Rasheed Taiwo Mudasiru, Ramatu Abubakar, Benjamin Ilunuamie Alenoghena
Quelle
Nature Journal of Emerging Sciences Technologies and Innovations
Publikation
2026-08-27
Band / Ausgabe
10 / 5
Seiten
533-548
ISSN / ISBN
3122-1017, 3115-4611
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Oluwaseun Adeniyi Ojerinde, Rasheed Taiwo Mudasiru, Ramatu Abubakar, Benjamin Ilunuamie Alenoghena (2026). <b>Deep learning-based multi-speaker separation and speech enhancement for forensic audio analysis</b>. Nature Journal of Emerging Sciences Technologies and Innovations, 10 (5), 533-548. https://doi.org/10.65752/fs075y71
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1