Can AI Win Gold? Assessing AI Math LLMs in K-12 International Competitions
Résumé fourni par la source
Large Language Models (LLMs) have shown impressive skills in various educational tasks. However, we have not fully explored how well they can solve K-12 international mathematics competition problems. In this paper, we aim to fill this gap by evaluating six advanced AI Math LLM tools: Doubao-1.5-Thinking-Pro-Vision, GPT-5, DeepSeek-R1, Claude-Opus-3, ChatGPT-4o-Latest, and Qwen3-VL-235B-A22B-T for their mathematical and algorithmic reasoning skills. we compiled 1230 Question-Answer from the K12 International Mathematics Competition representing various countries and regions, including 1074 competition questions (such as JMC, AMC8, JMO, IMO) and 156 areas of mathematics, namely Algebra, Geometry, Number Theory, Probability Theory, Combinatorics, and Logical Reasoning. According to the experimental results, Alibaba’s model Qwen3-VL-235B-A22B-T achieved the highest score 68.81% in all math topics. In contrast, the Number Theory column contains three upward arrows with data of 95.65%, 91.3%, and 65.22%, respectively, indicating three large models with the strongest ability in this topic. Similarly, the Probability topic shows three upward arrows with data of 87.5%, 65.3%, and 75%, respectively, demonstrating that large models also have a good understanding of this topic. In addition, Qwen3-vl-235b-a22b-t has obtained the best result for six times, respectively in AMC8, AMC10-b, JMC, AMC-level-C, AIME-2, BMO-1. The second is double-1.5-thinking-pro-vision, with five best scores in AMC12-B, AMC-level-D, AIME-1, JMO and IMO. In all models, AI Math LLMs did not get full marks, but human players can get full marks, which proves that the current human mathematical ability is greater than AI’s mathematical ability.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Can AI Win Gold? Assessing AI Math LLMs in K-12 International Competitions
- Date Crossref
- 21/11/2025
- Éditeur
- ACM
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.