Accès ouvert
2026
conference-paper
OpenAlex
Qingchen Yu, Shichao Song, Fang Ke, Zifan Zheng et autres
Existing benchmarks for Large Language Models (LLMs) are often static, knowledge-dependent, or costly, hindering the reliable evaluation of reasoning in dynamic user interactions. To address this, we propose TurtleBench, a novel benchmark derived from real user guesses in an online "Turtle Soup …
at, cn, us
(code pays fourni par la source)
2026
article
OpenAlex
Ziyu Liu, Yaping Huang, Shiqi Wang, Qingchen Yu
cn
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Zhiyu Li, Chunyan Xi, Chunyu Li, Shichao Song et autres
Large Language Models (LLMs) have become an essential infrastructure for Artificial General Intelligence (AGI), yet their lack of well-defined memory management systems hinders the development of long-context reasoning, continual personalization, and knowledge consistency.Existing models mainly rely on static parameters and short-lived contextual …
Accès ouvert
2025
article
OpenAlex
Qingchen Yu, Xin Liu, Qingguo Zhou, Chunming Wu
Recent advances in applying deep neural networks to programming tasks have achieved remarkable success in practice, prompting interest in exploring how well these models can perform traditional program analysis techniques. Data-flow analysis (DFA), a classic and well-established approach for analyzing programs, presents …
cn
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Qingchen Yu, Zifan Zheng, Chen Ding, Simin Niu et autres
The evaluation of large language models (LLMs) has traditionally relied on static benchmarks, a paradigm that poses two major limitations: (1) predefined test sets lack adaptability to diverse application domains, and (2) standardized evaluation protocols often fail to capture fine-grained assessments of …
Accès ouvert
2025
preprint
OpenAlex
Zhiyu Li, Shichao Song, Hanyu Wang, Simin Niu et autres
Large Language Models (LLMs) have emerged as foundational infrastructure in the pursuit of Artificial General Intelligence (AGI). Despite their remarkable capabilities in language perception and generation, current LLMs fundamentally lack a unified and structured architecture for handling memory. They primarily rely on …
Accès ouvert
2025
preprint
OpenAlex
Chen Ding, Qingchen Yu, Pengyuan Wang, Wentao Zhang et autres
With the release of OpenAI's o1 model, reasoning models that adopt slow-thinking strategies have become increasingly common. Their outputs often contain complex reasoning, intermediate steps, and self-reflection, making existing evaluation methods and reward models inadequate. In particular, they struggle to judge answer …
2025
conference-paper
OpenAlex
Qingchen Yu, Zifan Zheng, Chen Ding, Simin Niu et autres
2024
article
OpenAlex
Yue Zhang, Yining Wang, Yaping Huang, Qingchen Yu
cn
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Xun Liang, Shichao Song, Zifan Zheng, Hanyu Wang et autres
Large language models (LLMs) often exhibit deficient reasoning or generate hallucinations. To address these, studies prefixed with "Self-" such as Self-Consistency, Self-Improve, and Self-Refine have been initiated. They share a commonality: involving LLMs evaluating and updating themselves. Nonetheless, these efforts lack a …
Accès ouvert
2024
preprint
OpenAlex
Qingchen Yu, Zifan Zheng, Shichao Song, Zhiyu Li et autres
The continuous advancement of large language models (LLMs) has brought increasing attention to the critical issue of developing fair and reliable methods for evaluating their performance. Particularly, the emergence of cheating phenomena, such as test set leakage and prompt format overfitting, poses …
Accès ouvert
2024
preprint
OpenAlex
Ding Chen, Shichao Song, Qingchen Yu, Zhiyu Li et autres
Abstract In-context Learning (ICL) is one of the key methods for enhancing the performance of large language models on specific tasks by providing a set of few-shot question and answer examples. However, the ICL capability of different types of models shows significant …
cn, at
(code pays fourni par la source)