Improving Low-Resource Short Answer Scoring Through Large Language Model-Based Data Augmentation
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
The automated grading of subjective answers is crucial for reducing manual workload and enhancing feedback efficiency in online education, particularly for short answer scoring (SAS). However, in scenarios with limited labeled data, existing methods face challenges in sample diversity and scoring consistency due to limited training data and misaligned label distributions. While data augmentation and transfer learning have been employed to address these issues, rule-based approaches often lack textual variability, and general-domain embeddings struggle to align with domain-specific scoring criteria. To overcome these limitations, we propose the Scoring with Contextual Alignment and Language Enhancement (SCALE) framework, a novel LLM-driven training paradigm that synthesizes diverse responses while preserving scoring consistency. SCALE leverages a knowledge graph-based generation strategy to enhance sample diversity by substituting key phrases with contextually aligned alternatives and employs a style rewrite prompt to introduce linguistic variations. To mitigate label inconsistency, we introduce a polish align prompt that refines synthetic and real samples into a shared semantic subspace, training an annotator model for aligned scoring. Additionally, an entity-aware enhancement mechanism improves comprehension of formulas and quantitative content. Extensive experiments on multilingual and multi-domain datasets demonstrate that SCALE achieves state-of-the-art performance, improving Pearson scores by 4.9%, 2.27%, and 1.34% over BERT, RoBERTa, and ERNIE 3.0, respectively.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Improving Low-Resource Short Answer Scoring Through Large Language Model-Based Data Augmentation
- Date Crossref
- 01/05/2026
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.