Accès ouvert
2026
preprint
OpenAlex
Tianxin Wei, Zhan Shi, Minhua Lin, Bing He et autres
Learning from experience is critical for developing capable, self-improving large language model (LLM) agents. Existing methods typically extract knowledge from accumulated trajectories via reflection, memory, rules, or skills. However, agents in realistic environments continuously encounter novel tasks, often offering only a one-shot …
Accès ouvert
2026
preprint
OpenAlex
Ziyi Wang, Yuxuan Lu, Yimeng Zhang, Qun Liu et autres
Language model agents are increasingly effective in solving realistic tasks through multi-turn tool use. However, training reliable tool-using agents remains challenging in practice. While reinforcement learning provides an on-policy paradigm for improving agents from their own environment interactions, its effectiveness depends heavily …
Accès ouvert
2026
preprint
OpenAlex
Ziyi Wang, Yuxuan Lu, Yimeng Zhang, Qun Liu et autres
Language model agents are increasingly effective in solving realistic tasks through multi-turn tool use. However, training reliable tool-using agents remains challenging in practice. While reinforcement learning provides an on-policy paradigm for improving agents from their own environment interactions, its effectiveness depends heavily …
Accès ouvert
2026
preprint
OpenAlex
Zhan Shi, Bing He, Yisi Sang, Hüseyin Ozan Ҫirkinoğlu et autres
Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding what to keep. We report an autonomous system that runs this loop with no human in the loop, post-training a 30B Nemotron …
Accès ouvert
2026
preprint
OpenAlex
Zhan Shi, Bing He, Yisi Sang, Benoit Dumoulin et autres
Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding what to keep. We report an autonomous system that runs this loop with no human in the loop, post-training a 30B Nemotron …
us
(code pays fourni par la source)
Accès ouvert
2026
other
OpenAlex
Association for Computational Linguistics 2026, Bennett Bei, Yan Han, Qi He et autres
Recent research shows that LLM Agents can generate ``believable'' human behaviors via prompt-only methods, and such agents have been increasingly adopted in downstream applications. However, existing evaluation of these agents only focuses on qualitative believability (whether human raters think they are accurate), …
us, cn, mx
(code pays fourni par la source)
Accès ouvert
2026
other
OpenAlex
Association for Computational Linguistics 2026, Pei Chen, Ziwei Dong, Jiri Gesi et autres
Tool-calling agents are increasingly deployed in real-world customer-facing workflows. Yet most studies on tool-calling agents focus on idealized settings with general, fixed, and well-specified tasks. In real-world applications, user requests are often (1) ambiguous, (2) changing over time, or (3) infeasible due …
us, cn, mx
(code pays fourni par la source)
Accès ouvert
2026
other
OpenAlex
Association for Computational Linguistics 2026, Hansu Gu, Toby Li, Tun Lu et autres
Deploying machine learning models in real-world domain-specific scenarios is challenged by the scarcity of expert annotations and by data drift, where the statistical properties of incoming data continuously evolve. Active Learning (AL) iteratively improves compact models with expert annotations but suffers from …
us, cn, mx
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Zewen Liu, Zhan Shi, Yisi Sang, Bing He et autres
Auto-harness systems such as A-Evolve, GEPA, and Meta-Harness improve LLM agents by optimizing prompts, skills, tools, memories, and supporting infrastructure from execution feedback, but they are typically evaluated on fixed offline benchmarks. Real deployments instead present open-ended task streams: histories grow without …
Accès ouvert
2026
preprint
OpenAlex
Zewen Liu, Zhan Shi, Yisi Sang, Bing He et autres
Auto-harness systems such as A-Evolve, GEPA, and Meta-Harness improve LLM agents by optimizing prompts, skills, tools, memories, and supporting infrastructure from execution feedback, but they are typically evaluated on fixed offline benchmarks. Real deployments instead present open-ended task streams: histories grow without …
us, de, jm, mx
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Minhua Lin, Juncheng Wu, Zijun Wang, Zhan Shi et autres
LLM agents are increasingly deployed as systems built around editable external harnesses, including prompts, skills, memories and tools, that shape task execution without changing model parameters. Harness self-evolution adapts such agents by updating these harnesses from execution evidence. Yet it remains unclear …
Accès ouvert
2026
preprint
OpenAlex
Minhua Lin, Juncheng Wu, Zijun Wang, Zhan Shi et autres
LLM agents are increasingly deployed as systems built around editable external harnesses, including prompts, skills, memories and tools, that shape task execution without changing model parameters. Harness self-evolution adapts such agents by updating these harnesses from execution evidence. Yet it remains unclear …
us, de, jm, mx
(code pays fourni par la source)