Speech Emotion Recognition Using Hybrid Deep Learning Models and Diverse Acoustic Features
Rattachement africain : cy. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Deep learning models have recently gained popularity in audio classification tasks due to their exceptional efficiency and accuracy. With the help of deep learning models, complex patterns can be learned and trained. In this work, we tested performance of multiple deep learning models for speech emotion recognition, including a hybrid CNN-LSTM model, by utilizing diverse acoustic input features. In particular, raw audio features, as well as, three types of spectrograms, namely Logarithmic Frequency Scale Spectrogram, Mel-frequency Cepstral Coefficients (MFCC) and Short Time Fourier Transform (STFT) were utilized as input features. Then, these raw audio data and different types of spectrograms were input to various deep learning models for speech emotion classification: Artificial Neural Network (ANN), Long Short-Term Memory Network (LSTM), Convolutional Neural Network (CNN) and CNN-LSTM (hybrid model). The primary objective is to develop a robust system that can accurately identify, classify and predict human emotions from audio recordings. In addition, performances of different audio features were tested on various deep learning models. The CREMA-D dataset was employed to assess the performance of that is consisting of six classes of human emotions (Anger, Neutral, Disgust, Sad, Happy and Fear). On the raw audio data, the ANN model performed better than the remaining models. When Logarithmic Frequency Scale Spectrogram features are used by CNN and CNN-LSTM models, the best results were obtained with accuracies of 99.80% and 100% respectively. Our experiments indicate that for emotion recognition on the CREMA-D dataset, Logarithmic Frequency Scale Spectrogram achieved the best and more consistent results using various deep learning models.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Speech Emotion Recognition Using Hybrid Deep Learning Models and Diverse Acoustic Features
- Date Crossref
- 23/05/2025
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.