Aller au contenu principal
Accès ouvert déclaré 2026 dataset

DialogPII - Multilingual Synthetic Dialogs and Transcripts Annotated for Personally Identifiable Information

0Citations signalées, ce qui n’est pas une note de qualité
7Institutions déclarées
3Pays d’affiliation déclarés

Rattachement africain : de, nl, ar. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

DialogPII is a multilingual synthetic dialogue dataset created for research on conversational de-identification and the detection of personally identifiable information (PII) in dialogue transcripts. The dataset contains synthetic dialog transcripts in 11 languages across 8 interaction scenarios, including emergencycalls, medical anamnesis interviews, therapy and group therapy sessions, insurance communication, customer support, police reports, and interviews regarding an AI-supported dashboard used in a medical context. This dataset is designed to support the development and evaluation of privacy-preserving NLP methods for conversational data. Unlike many document-based de-identification resources, this dataset focuses on dialog settings, where personal information may be fragmented across turns, expressed informally, or embedded in context-dependent utterances. The annotations cover direct and contextual PII categories, including names, social relations, email addresses, phone numbers, URLs, locations, organizations, professions, products, dates and times, ages, quantities, and miscellaneous identifying information. The dataset can be used for tasks such as multilingual PII detection, dialog de-identification, and the development of privacy-aware NLP pipelines. This release includes DialogPII: the annotated original dialogs and transcripts in both .json (span offsets) and .txt (inline tags) files, the annotation scheme and guidelines used to create the dataset, and reviewer guidelines for the analysis of the automatically translated texts. Our model finetuned on DialogPII can be found here: https://huggingface.co/DFKI-SLT/multilingual_DialogPII_NER This Zenodo record serves as an external resource for the paper “DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information”. Please cite the associated paper and this Zenodo record when using the dataset.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.