Aller au contenu principal
Accès ouvert déclaré 2026 article

Speech-based early depressive symptom screening using MAX-Net: a multi-scale autoencoder-augmented xLSTM framework

0Citations signalées — pas une note de qualité
4Institutions déclarées
2Pays d’affiliation déclarés

Résumé fourni par la source

Background Early screening for depressive symptoms lacks objective and stable biomarkers, while traditional methods based on questionnaires and subjective assessments suffer from issues such as low efficiency. Although the use of artificial intelligence to analyze speech signals has become a growing trend, existing methods are limited by challenges such as small medical datasets, complex features, and long-term temporal dependencies, resulting in insufficient recognition accuracy and poor cross-dataset generalization performance. Methods To address the above issues, this paper proposes MAX-Net, a multi-scale feature-enhanced long-term sequential autoencoder network based on an improved extended long short-term memory (xLSTM) for depressive symptom screening. MAX-Net combines the improved xLSTM with autoencoder feature enhancement and introduces a Multi-Scale Convolutional layer (MSC) module. To mitigate the issue of data sparsity, we propose an autoencoder-based feature space enhancement strategy. Furthermore, we investigate the impact of varying numbers of augmented samples on model performance. The model is trained and evaluated on the in-house dataset HMU-DDAC-24 and retrained on the public dataset MODMA using the same architecture and hyperparameters to assess its stability and cross-dataset applicability. Results Experiments on HMU-DDAC-24 show that MAX-Net achieves an Area Under the ROC Curve (AUC) of 0.945 and a Mean Absolute Error (MAE) of 1.61 in the depression symptom classification task, significantly outperforming existing baseline models. It maintains competitive performance on the MODMA dataset (AUC = 0.804), suggesting architectural stability and consistent effectiveness across different data sources when retrained. Further experiments revealed that the optimal data augmentation strategy varies across different datasets: on HMU-DDAC-24, performance peaks when 15,000 samples are generated, whereas on MODMA, performance stabilizes when 18,000–22,000 samples are generated. Conclusion MAX-Net provides an effective framework for speech-based early depression screening and demonstrates consistent performance across datasets under a unified configuration. These findings indicate that the proposed architecture is robust to variations in data distribution. Furthermore, the autoencoder-based augmentation strategy has proven beneficial in low-data-volume scenarios. Overall, this study supports the application potential of MAX-Net as an auxiliary tool for early depression screening.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Speech-based early depressive symptom screening using MAX-Net: a multi-scale autoencoder-augmented xLSTM framework
Date Crossref
03/09/2026
Éditeur
Frontiers Media SA
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Mental Health via WritingEmotion and Mood RecognitionDigital Mental Health Interventions

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.