Aller au contenu principal
Accès ouvert déclaré 2026 article

Investigating the Impact of Combined Spectral and Prosodic Features on Speech Emotion Recognition

3Citations signalées, ce qui n’est pas une note de qualité
4Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : sa, pk. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Emotions play a fundamental role in human cognition, behavior, and social interaction, making automatic recognition a key topic in affective computing. Many existing approaches to recognizing emotions from speech rely heavily on Mel-Frequency Cepstral Coefficients (MFCCs), which capture short-term spectral features but insufficiently represent prosody, long-term dynamics, and tonal nuances that are critical for accurate classification. This study presents an interpretable and computationally efficient framework for recognizing emotions from speech by employing a compact set of fourteen spectral and prosodic acoustic features, including pitch, shimmer, jitter, loudness, harmonic-to-noise ratio (HNR), and measures of temporal variation. Using tree-based ensemble methods, the proposed system achieved its best performance with the XGBoost classifier, reaching an accuracy of 96.79% on the Toronto Emotional Speech Set (TESS). Statistical validation using the Kruskal–Wallis test and effect size analysis revealed that HNR, mean pitch, and shimmer were the most discriminative predictors of emotional state, thereby providing transparency into the classification process. The system also demonstrated real-time capability with inference times between 201 and 231 milliseconds, confirming that accurate, efficient, and interpretable speech emotion recognition can be achieved without relying solely on deep learning models.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Investigating the Impact of Combined Spectral and Prosodic Features on Speech Emotion Recognition
Date Crossref
01/01/2026
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Emotion and Mood RecognitionMusic and Audio ProcessingSpeech Recognition and Synthesis

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.