Dataset of anonymized discharge summaries of sepsis patients from a Brazilian tertiary hospital for NLP applications.
Résumé fourni par la source
The availability of Brazilian Portuguese health record text datasets for Natural Language Processing (NLP) applications is limited, especially for educational purposes. The main reason for this is data sensitivity, which dictates the need for accurate data anonymization. This article describes a new dataset compiled to help bridge the gap in publicly available information in this area. The data were extracted from discharge summaries in the electronic health record system of a tertiary teaching hospital (Hospital das Clínicas, Ribeirão Preto Medical School). Health records were filtered to identify adult patients diagnosed with sepsis (ICD-10). This diagnosis was chosen because patients with sepsis generally require an extended stay in the hospital and, therefore, have discharged summaries with long text. The data were curated manually by a physician to exclude discharge summaries with incomplete descriptions, resulting in 387 cases. The texts were processed to exclude special characters, expand standard medical abbreviations and anonymize the data. The anonymization process was conducted in two steps: unsupervised anonymization using GLiNER, followed by supervised anonymization using a spaCy model trained by the author to identify named entities. Data related to key structured clinical variables (length of hospital stay, number of specialties involved, ICU admission, palliative care status, discharge outcome) were also extracted from the original health records and combined with each summary. A manual medical record review of the selected cases was performed to ensure data quality and the efficacy of anonymization, excluding all records that did not contain relevant medical information. The resulting dataset comprises 200 anonymized Brazilian Portuguese discharge summaries, along with their respective associated variables, displayed in a tabular format. The data provided in this article offers a valuable, practical resource featuring real medical data for teaching and learning basic NLP techniques, such as text preprocessing, named entity recognition, text classification, and topic modeling.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Dataset of anonymized discharge summaries of sepsis patients from a Brazilian tertiary hospital for NLP applications.
- Date Crossref
- 01/08/2025
- Éditeur
- Elsevier BV
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.