SINAI at SemEval-2024 Task 8: Fine-tuning on Words and Perplexity as Features for Detecting Machine Written Text
Rattachement africain : es. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
This work describes the system submitted by the SINAI team to the subtask A of Task 8 of SemEval 2024, as well as two additional systems evaluated during the training phase of the shared task.We claim that the perplexity score of a text may be used as a classification signal.Accordingly, we conduct a study on the utility of perplexity for discerning text authorship, and we perform a comparative analysis of the results obtained on the datasets of the task.The results of this study motivated us to use as classification features the word embeddings vectors of the input texts and its corresponding perplexity score.Likewise, the submitted system is a fine-tuning version of the XLM-RoBERTa-Large model.The analysis of the results of the evaluation shows large differences among the language probability distribution of the training and test sets.Nonetheless, the results show that perplexity can be used as feature for identifying machine generated text, hence our claim holds.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- SINAI at SemEval-2024 Task 8: Fine-tuning on Words and Perplexity as Features for Detecting Machine Written Text
- Date Crossref
- 01/01/2024
- Éditeur
- Association for Computational Linguistics
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.