Aller au contenu principal
Accès ouvert déclaré 2025 conference-paper

Can AI Win Gold? Assessing AI Math LLMs in K-12 International Competitions

0Citations signalées — pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Large Language Models (LLMs) have shown impressive skills in various educational tasks. However, we have not fully explored how well they can solve K-12 international mathematics competition problems. In this paper, we aim to fill this gap by evaluating six advanced AI Math LLM tools: Doubao-1.5-Thinking-Pro-Vision, GPT-5, DeepSeek-R1, Claude-Opus-3, ChatGPT-4o-Latest, and Qwen3-VL-235B-A22B-T for their mathematical and algorithmic reasoning skills. we compiled 1230 Question-Answer from the K12 International Mathematics Competition representing various countries and regions, including 1074 competition questions (such as JMC, AMC8, JMO, IMO) and 156 areas of mathematics, namely Algebra, Geometry, Number Theory, Probability Theory, Combinatorics, and Logical Reasoning. According to the experimental results, Alibaba’s model Qwen3-VL-235B-A22B-T achieved the highest score 68.81% in all math topics. In contrast, the Number Theory column contains three upward arrows with data of 95.65%, 91.3%, and 65.22%, respectively, indicating three large models with the strongest ability in this topic. Similarly, the Probability topic shows three upward arrows with data of 87.5%, 65.3%, and 75%, respectively, demonstrating that large models also have a good understanding of this topic. In addition, Qwen3-vl-235b-a22b-t has obtained the best result for six times, respectively in AMC8, AMC10-b, JMC, AMC-level-C, AIME-2, BMO-1. The second is double-1.5-thinking-pro-vision, with five best scores in AMC12-B, AMC-level-D, AIME-1, JMO and IMO. In all models, AI Math LLMs did not get full marks, but human players can get full marks, which proves that the current human mathematical ability is greater than AI’s mathematical ability.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Can AI Win Gold? Assessing AI Math LLMs in K-12 International Competitions
Date Crossref
21/11/2025
Éditeur
ACM
Type
proceedings-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Intelligent Tutoring Systems and Adaptive LearningMathematics, Computing, and Information ProcessingMathematics Education and Teaching Techniques

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, ROR et la Banque mondiale, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune donnée externe enregistrée en base. Sources et limites.