Aller au contenu principal
2024 article

Enhanced Speech Processing for Air Traffic Control Using an Optimized Generative Adversarial Network

1Citations signalées, ce qui n’est pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés

Le résumé fourni par la source

In order to improve the quality of voice calls during air traffic control, an improved SEGAN air traffic control speech enhancement algorithm is proposed. Aiming at the problem that the traditional speech enhancement algorithm based on generative adversarial network (SEGAN) is drowned under low signal-to-noise ratio conditions, a multi-stage, multi-mapping, multi-dimensional output generator and a multi-scale, multi-discriminator network model are proposed based on the SEGAN network model. First, the speech semantic features are extracted based on the deep neural network structure to complete the semantic segmentation of air traffic control speech. Secondly, multiple generators are set to further optimize the speech signal. Then, a downsampling module is added to the convolution layer to improve the model's utilization of speech information and reduce the loss of speech information. Finally, multi-scale, multiple discriminators are used to learn the distribution law and information of speech samples in multiple directions. The results show that under low signal-to-noise ratio conditions, the improved SEGAN model improves the short-term objective intelligibility (STOI) and the perceptual evaluation of speech quality (PESQ) by and respectively , which can quickly and effectively perform air traffic control speech enhancement and provide preparation for subsequent air traffic control speech recognition.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Enhanced Speech Processing for Air Traffic Control Using an Optimized Generative Adversarial Network
Date Crossref
18/10/2024
Éditeur
Cresta Press
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les sujets associés

Speech and Audio ProcessingSpeech Recognition and SynthesisAerodynamics and Acoustics in Jet Flows

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.