Aller au contenu principal
Accès ouvert déclaré 2026 article

Benchmarking large language models for target-specific financial stance detection in 10-K MD&A sections and earnings call transcripts

0Citations signalées — pas une note de qualité
3Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Financial disclosures contain rich narrative information about firm performance, but their length, specialized terminology, and target-dependent language make sentence-level analysis challenging. Existing financial sentiment methods typically assess overall positive or negative tone, rather than stance toward specific financial targets. In this study, we introduce a sentence-level benchmark for target-specific financial stance detection in Form 10-K Management’s Discussion and Analysis (MD&A) sections and quarterly earnings call transcripts (ECTs). The benchmark focuses on three financial targets–debt, earnings per share (EPS), and sales–and assigns each target-relevant sentence a positive, negative, or neutral stance label. The corpus is drawn from five public companies. The study is intended as a benchmark for NLP methods and does not support broad claims about financial disclosures in general due to the limited set of companies and financial targets. For scalability, the training split is labeled using ChatGPT-o3-pro, while the held-out test split is independently annotated by human annotators and adjudicated to form a human-consensus gold standard. Using this benchmark, we evaluate four contemporary large language models under zero-shot, few-shot, Chain-of-Thought, and document-context prompting conditions. Model outputs are assessed using accuracy, macro-averaged and weighted precision, recall, and F1, with paired statistical tests used to evaluate performance differences. Results show that LLMs can provide useful baselines for low-label, target-specific financial stance detection, but performance varies across models, document types, targets, and prompting strategies. GPT-4.1-mini and Gemma4-31B generally achieved the strongest overall performance, followed by Llama 3.3-70B and Mistral Small 3.2-24B, although no single model or prompting strategy dominated across all settings. The no-document-context condition achieved the highest Macro-F1 score. However, this result is best interpreted as alignment with the sentence-level annotation protocol, in which annotators did not have access to the full document context, rather than as evidence that document context is generally unhelpful. These findings highlight both the promise and the remaining limitations of LLM-based financial stance detection, particularly the need for broader industry coverage and economic validation.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé, mais le titre doit être comparé manuellement.

Titre Crossref
Benchmarking large language models for target-specific financial stance detection in 10-K MD&A sections and earnings call transcripts
Date Crossref
01/09/2026
Éditeur
Springer Science and Business Media LLC
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Auditing, Earnings Management, GovernanceArtificial Intelligence in Healthcare and EducationText Readability and Simplification

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.