IJAEMS
This paper introduces a new framework for the identification of artificial (Deep Fake Audio) voice using both the feature level and image level learning approaches in deep learning based technique. The rapid development in Artificial Intelligence and generation techniques has made the creation of highly believable audio possible; this poses serious threats to digital security and credibility of the information. In order to deal with this issue, the proposed system consists of a two-branch system consisting of Mel Frequency Cepstral Coefficient (MFCC) and Gammatone Frequency Cepstral Coefficient (GFCC) with a feature level approach using a Long Short-Term Memory (LSTM) model and Mel spectrogram images with an image-level approach using Deep Convolutional Neural Network (DCNN). The system uses advanced techniques of preprocessing, feature extraction and CatBoost based feature fusion. Experiments have been performed using the ASV spoof 2021 dataset for testing the efficacy of the proposed system in discriminating between real and fake audio. The system gives better accuracy, precision and robustness as compared to some existing state-of-the-art techniques. The system can also be used for real-time detection through graphical user interface.