International Journal of Advanced Engineering and Management System Logo IJAEMS
Back
Research Article Matrix Node

Deep Learning Based Feature Level and Image Level Framework for Fake voice Detection

Author(s): S. Sharanyaa, Shalini S, Madhumitha P, Sherlin Jemima J
Volume RegistryVolume 2
Issue PeriodIssue 03
Published Date08 Aug 2026
DOI Identifier

Abstract

This paper introduces a new framework for the identification of artificial (Deep Fake Audio) voice using both the feature level and image level learning approaches in deep learning based technique. The rapid development in Artificial Intelligence and generation techniques has made the creation of highly believable audio possible; this poses serious threats to digital security and credibility of the information. In order to deal with this issue, the proposed system consists of a two-branch system consisting of Mel Frequency Cepstral Coefficient (MFCC) and Gammatone Frequency Cepstral Coefficient (GFCC) with a feature level approach using a Long Short-Term Memory (LSTM) model and Mel spectrogram images with an image-level approach using Deep Convolutional Neural Network (DCNN). The system uses advanced techniques of preprocessing, feature extraction and CatBoost based feature fusion. Experiments have been performed using the ASV spoof 2021 dataset for testing the efficacy of the proposed system in discriminating between real and fake audio. The system gives better accuracy, precision and robustness as compared to some existing state-of-the-art techniques. The system can also be used for real-time detection through graphical user interface.

Keywords

Deepfake Audio MFCC GFCC Long Short Term memory Convolutional Neural Network Deep Learning Spectrogram

References (22)

  1. [Multimodaltrace: Deepfake Detection using Audiovisual Representation Learning,Muhammad Anas Raza Khalid Mahmood Malik,Oakland University,2023 IEEE/CVF Conference on Computer Vision and Pat-tern Recognition Workshops (CVPRW) — 979-83503- 0249-3/23/31.00©2023IEEE—DOI: https://doi.org/10.1109/CVPRW59228.2023.00106
  2. C. Stupp, “Fraudsters used Ai to mimic CEO’s voice in unusual cybercrime case,” Wall Street J., vol. 30, no. 8, pp. 12, 2019.
  3. Deepfake audio detection by speaker verification,Alessandro Pianese,Davide Cozzolino, Giovanni Poggi and Luisa Verdoliva ,University Federico II of Naples, Italy,2022 IEEE International Workshop on Information Forensics and Security (WIFS) — 979- 83503-0967-6/22/31.00©2022IEEE—DOI: https://doi.org/10.1109/WIFS55849.2022.9975428
  4. Deepfake detection using deep learning methods: A systematic and comprehensive review , Arash Heidari, Nima Jafari Navimipour, Hasan Dag, Mehmet Unal, 2023,https://doi.org/10.1002/widm.1520
  5. D. Garg and R. Gill, ”Deepfake Generation and Detection - An Exploratory Study,” 2023 10th IEEE Uttar Pradesh Section International Conference on Electrical, Electronics and Computer Engineering (UPCON), Gautam Buddha Nagar, India, 2023, pp. 888 893, doi: https://doi.org/10.1109/UPCON59197.2023.10434896

Format Citation Record

S. Sharanyaa, Shalini S, Madhumitha P, and Sherlin Jemima J, "Deep Learning Based Feature Level and Image Level Framework for Fake voice Detection," Int. J. Adv. Eng. Manag. Syst., vol. 2, no. 3, pp. 380-387, 2026.