Aller au contenu principal
Accès ouvert déclaré 2025 article

Do transformer-based token classification methods solve the problem of terminology extraction?

0Citations signalées, ce qui n’est pas une note de qualité
3Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : pl, cz. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Abstract Results obtained by transformer-based token classification models are now considered to be a benchmark for the Automatic Terminology Extraction (ATE) task. However, the unsatisfactory results (they rarely exceed 0.7 of the F1 value) raise the question of whether this approach is correct and of what text features are being remembered or inferred by the model trained on this type of annotation. In the paper, we describe a number of experiments using the fine-tuned RoBERTa base model on the ACTER data, RD-TEC, and three Wikipedia articles, which proved that the results of the ATE task obtained by such models depend considerably on the type of texts being processed and their relationship to the training data. While the results are relatively good for some texts with highly specialized vocabulary, the poor results seem to correlate with the high frequency (in general English texts) of tokens that are part of terms in a particular domain. Another property that affects the results is the degree of overlap between the vocabulary of the test data and the vocabulary of terms from the training data. Words that have been labeled as terms in the training data are usually labeled as terms in other, unrelated domains as well. Moreover, we show that the results obtained by these models are unstable—models trained on more data do not include all the items identified by models trained on a smaller dataset and can present substantially lower performance.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Do transformer-based token classification methods solve the problem of terminology extraction?
Date Crossref
15/08/2025
Éditeur
Cambridge University Press (CUP)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Institute of Computer Science pays non établi dans la notice
    Structure de recherche
  • Czech Academy of Sciences pays non établi dans la notice
    Structure de recherche
  • Polish Academy of Sciences pays non établi dans la notice
    Organisme public

Institute of Computer Science, Czech Academy of Sciences et Polish Academy of Sciences.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Advanced Text Analysis TechniquesBiomedical Text Mining and OntologiesSemantic Web and Ontologies

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.