Aller au contenu principal
2026 article

Improving Low-Resource Short Answer Scoring Through Large Language Model-Based Data Augmentation

1Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

The automated grading of subjective answers is crucial for reducing manual workload and enhancing feedback efficiency in online education, particularly for short answer scoring (SAS). However, in scenarios with limited labeled data, existing methods face challenges in sample diversity and scoring consistency due to limited training data and misaligned label distributions. While data augmentation and transfer learning have been employed to address these issues, rule-based approaches often lack textual variability, and general-domain embeddings struggle to align with domain-specific scoring criteria. To overcome these limitations, we propose the Scoring with Contextual Alignment and Language Enhancement (SCALE) framework, a novel LLM-driven training paradigm that synthesizes diverse responses while preserving scoring consistency. SCALE leverages a knowledge graph-based generation strategy to enhance sample diversity by substituting key phrases with contextually aligned alternatives and employs a style rewrite prompt to introduce linguistic variations. To mitigate label inconsistency, we introduce a polish align prompt that refines synthetic and real samples into a shared semantic subspace, training an annotator model for aligned scoring. Additionally, an entity-aware enhancement mechanism improves comprehension of formulas and quantitative content. Extensive experiments on multilingual and multi-domain datasets demonstrate that SCALE achieves state-of-the-art performance, improving Pearson scores by 4.9%, 2.27%, and 1.34% over BERT, RoBERTa, and ERNIE 3.0, respectively.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Improving Low-Resource Short Answer Scoring Through Large Language Model-Based Data Augmentation
Date Crossref
01/05/2026
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Topic ModelingExpert finding and Q&A systemsInformation Retrieval and Search Behavior

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.