Aller au contenu principal
Accès ouvert déclaré 2025 preprint

Novel document-level measures and their performance for active learning uncertainty sampling—a use case for automatic CDSS ontology curation

1Citations signalées, ce qui n’est pas une note de qualité
11Institutions déclarées
4Pays d’affiliation déclarés

Rattachement africain : us, gb, dk, no. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Objective: To explore new strategies to make the document selection process more transparent, reproducible, and effective for the active learning process. The ultimate goal is to leverage active learning in identifying keyphrases to facilitate ontology development and construction, to streamline the process, and help with the long-term maintenance. Methods: The active learning pipeline used a BILSTM-CRF model and over 2900 abstracts retrieved from PubMed relevant to clinical decision support systems. We started the model training with synthetic labeled abstracts, then used different strategies to select domain experts' annotated abstracts (gold standards). Random sampling was used as the baseline. Recall, F1 (beta = 1, 5, and 10) scores are used as measures to compare the performance of the active learning pipeline by different strategies. Results: ) for recall and F1. The systematic evaluations show that KPSum (actual order) shows consistent improvement in both recall and F1 and KPSum (actual order) shows better results than the random sampling results. The document order (actual versus reverse) does not seem to play a critical role across strategies in model learning and performance in our datasets, although in some strategies, actual order shows slightly more effective results. The weighted F1 (beta = 5 and 10) provided complementary results to raw recall and F1 (beta = 1). Conclusion: While prior work on uncertainty sampling typically focuses on token-level uncertainty metrics within generic NER tasks, our work advances this line of research by introducing a higher-level abstraction: document-level uncertainty aggregation. With a human-in-the-loop Active Learning pipeline, it can effectively prioritize high-impact documents, improve early-cycle recall, and reduce annotation effort. Our results show promise in automating part of ontology construction and maintenance work, i.e., monitoring and screening new publications to identify candidate keyphrases. However, future work needs to improve the model performance to make it usable in real-world operations.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Novel document-level measures and their performance for active learning uncertainty sampling—a use case for automatic CDSS ontology curation
Date Crossref
17/04/2025
Éditeur
openRxiv
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Clemson University pays non établi dans la notice
    Université ou école supérieure
  • University of Pittsburgh pays non établi dans la notice
    Université ou école supérieure
  • George Mason University pays non établi dans la notice
    Université ou école supérieure
  • The University of Texas Health Science Center pays non établi dans la notice
    Université ou école supérieure
  • The University of Texas Health Science Center at Houston pays non établi dans la notice
    Université ou école supérieure
  • University of Cumbria pays non établi dans la notice
    Université ou école supérieure
  • Indiana University – Purdue University Indianapolis pays non établi dans la notice
    Université ou école supérieure
  • Vanderbilt University Medical Center pays non établi dans la notice
    Établissement de santé
  • Aalborg University pays non établi dans la notice
    Université ou école supérieure
  • Ohio University Ohio Musculoskeletal and Neurologic Institute pays non établi dans la notice
    Université ou école supérieure
  • Norwegian University of Science and Technology pays non établi dans la notice
    Université ou école supérieure
  • School of Computing pays non établi dans la notice
    Université ou école supérieure

Clemson University, University of Pittsburgh et George Mason University, avec 9 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Biomedical Text Mining and OntologiesAdvanced Text Analysis TechniquesTopic Modeling

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.