A Secure Multilingual Knowledge Extraction Pipeline for Virtual Collaboration Using Generative AI
Résumé fourni par la source
The move to working from home and mixing office work with homework has caused meetings where people speak languages. These meetings have transcripts that include company details. These transcripts are usually not recorded, summarized or stored safely so they can be used in languages. Previous research focused on making summaries translating, keeping information secure and following GDPR rules. There was no solution. This paper describes the design and testing of a system that finds and removes information identifies who spoke makes summaries and translates all while keeping GDPR rules in mind inside a Streamlit application. When I tested the system on a transcript with speakers, I got speaker detection with a score of 1.000 and clear summaries in all languages. The evaluation scores were ROUGE-1 0.386 ROUGE-2 0.207 ROUGE-L 0.341 and BLEU 7.07. However, some evaluation metrics were not correct: the average F1 score was, outside the expected range. The historical dashboard showed NaN values for precision and recall. These problems mean that even though the system works for language processing and functions the security and metric validation parts still need to be improved before the system can claim to be fully compliant.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.