Do transformer-based token classification methods solve the problem of terminology extraction?
Rattachement africain : pl, cz. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Abstract Results obtained by transformer-based token classification models are now considered to be a benchmark for the Automatic Terminology Extraction (ATE) task. However, the unsatisfactory results (they rarely exceed 0.7 of the F1 value) raise the question of whether this approach is correct and of what text features are being remembered or inferred by the model trained on this type of annotation. In the paper, we describe a number of experiments using the fine-tuned RoBERTa base model on the ACTER data, RD-TEC, and three Wikipedia articles, which proved that the results of the ATE task obtained by such models depend considerably on the type of texts being processed and their relationship to the training data. While the results are relatively good for some texts with highly specialized vocabulary, the poor results seem to correlate with the high frequency (in general English texts) of tokens that are part of terms in a particular domain. Another property that affects the results is the degree of overlap between the vocabulary of the test data and the vocabulary of terms from the training data. Words that have been labeled as terms in the training data are usually labeled as terms in other, unrelated domains as well. Moreover, we show that the results obtained by these models are unstable—models trained on more data do not include all the items identified by models trained on a smaller dataset and can present substantially lower performance.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Do transformer-based token classification methods solve the problem of terminology extraction?
- Date Crossref
- 15/08/2025
- Éditeur
- Cambridge University Press (CUP)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Institute of Computer Science pays non établi dans la noticeStructure de recherche
-
Czech Academy of Sciences pays non établi dans la noticeStructure de recherche
-
Polish Academy of Sciences pays non établi dans la noticeOrganisme public
Institute of Computer Science, Czech Academy of Sciences et Polish Academy of Sciences.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.