DialogPII - Multilingual Synthetic Dialogs and Transcripts Annotated for Personally Identifiable Information
Rattachement africain : de, nl, ar. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
DialogPII is a multilingual synthetic dialogue dataset created for research on conversational de-identification and the detection of personally identifiable information (PII) in dialogue transcripts. The dataset contains synthetic dialog transcripts in 11 languages across 8 interaction scenarios, including emergencycalls, medical anamnesis interviews, therapy and group therapy sessions, insurance communication, customer support, police reports, and interviews regarding an AI-supported dashboard used in a medical context. This dataset is designed to support the development and evaluation of privacy-preserving NLP methods for conversational data. Unlike many document-based de-identification resources, this dataset focuses on dialog settings, where personal information may be fragmented across turns, expressed informally, or embedded in context-dependent utterances. The annotations cover direct and contextual PII categories, including names, social relations, email addresses, phone numbers, URLs, locations, organizations, professions, products, dates and times, ages, quantities, and miscellaneous identifying information. The dataset can be used for tasks such as multilingual PII detection, dialog de-identification, and the development of privacy-aware NLP pipelines. This release includes DialogPII: the annotated original dialogs and transcripts in both .json (span offsets) and .txt (inline tags) files, the annotation scheme and guidelines used to create the dataset, and reviewer guidelines for the analysis of the automatically translated texts. Our model finetuned on DialogPII can be found here: https://huggingface.co/DFKI-SLT/multilingual_DialogPII_NER This Zenodo record serves as an external resource for the paper “DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information”. Please cite the associated paper and this Zenodo record when using the dataset.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.