Aller au contenu principal
Profil bibliographique

Qingchen Yu

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

20Publications signalées
92Citations signalées
2Affiliations récentes

Les institutions déclarées

Les domaines associés

Natural Language Processing TechniquesTopic ModelingAdvanced Malware Detection TechniquesSemantic Web and OntologiesSoftware Engineering Research

Les publications récentes

Accès ouvert 2026 conference-paper OpenAlex

TURTLEBENCH: Evaluating Top Language Models via Real-World Yes/No Puzzles

Qingchen Yu, Shichao Song, Fang Ke, Zifan Zheng et autres

Existing benchmarks for Large Language Models (LLMs) are often static, knowledge-dependent, or costly, hindering the reliable evaluation of reasoning in dynamic user interactions. To address this, we propose TurtleBench, a novel benchmark derived from real user guesses in an online "Turtle Soup …

at, cn, us (code pays fourni par la source)

0 citations
Accès ouvert 2025 preprint OpenAlex

MemOS: A Memory OS for AI System

Zhiyu Li, Chunyan Xi, Chunyu Li, Shichao Song et autres

Large Language Models (LLMs) have become an essential infrastructure for Artificial General Intelligence (AGI), yet their lack of well-defined memory management systems hinders the development of long-context reasoning, continual personalization, and knowledge consistency.Existing models mainly rely on static parameters and short-lived contextual …

0 citations arXiv (Cornell University)
Accès ouvert 2025 article OpenAlex

Bridging the Gaps between Graph Neural Networks and Data-Flow Analysis: The Closer, the Better

Qingchen Yu, Xin Liu, Qingguo Zhou, Chunming Wu

Recent advances in applying deep neural networks to programming tasks have achieved remarkable success in practice, prompting interest in exploring how well these models can perform traditional program analysis techniques. Data-flow analysis (DFA), a classic and well-established approach for analyzing programs, presents …

cn (code pays fourni par la source)

0 citations Proceedings of the ACM on software engineering.
Accès ouvert 2025 preprint OpenAlex

GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning

Qingchen Yu, Zifan Zheng, Chen Ding, Simin Niu et autres

The evaluation of large language models (LLMs) has traditionally relied on static benchmarks, a paradigm that poses two major limitations: (1) predefined test sets lack adaptability to diverse application domains, and (2) standardized evaluation protocols often fail to capture fine-grained assessments of …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models

Zhiyu Li, Shichao Song, Hanyu Wang, Simin Niu et autres

Large Language Models (LLMs) have emerged as foundational infrastructure in the pursuit of Artificial General Intelligence (AGI). Despite their remarkable capabilities in language perception and generation, current LLMs fundamentally lack a unified and structured architecture for handling memory. They primarily rely on …

2 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

xVerify: Efficient Answer Verifier for Reasoning Model Evaluations

Chen Ding, Qingchen Yu, Pengyuan Wang, Wentao Zhang et autres

With the release of OpenAI's o1 model, reasoning models that adopt slow-thinking strategies have become increasingly common. Their outputs often contain complex reasoning, intermediate steps, and self-reflection, making existing evaluation methods and reward models inadequate. In particular, they struggle to judge answer …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

Internal Consistency and Self-Feedback in Large Language Models: A Survey

Xun Liang, Shichao Song, Zifan Zheng, Hanyu Wang et autres

Large language models (LLMs) often exhibit deficient reasoning or generate hallucinations. To address these, studies prefixed with "Self-" such as Self-Consistency, Self-Improve, and Self-Refine have been initiated. They share a commonality: involving LLMs evaluating and updating themselves. Nonetheless, these efforts lack a …

19 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation

Qingchen Yu, Zifan Zheng, Shichao Song, Zhiyu Li et autres

The continuous advancement of large language models (LLMs) has brought increasing attention to the critical issue of developing fair and reliable methods for evaluating their performance. Particularly, the emergence of cheating phenomena, such as test set leakage and prompt format overfitting, poses …

2 citations arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.