Accès ouvert
2026
preprint
OpenAlex
John Yang, K. Lieret, Jeffrey Ma, Parth Thakkar et autres
Turning ideas into full software projects from scratch has become a popular use case for language models. Agents are being deployed to seed, maintain, and grow codebases over extended periods with minimal human oversight. Such settings require models to make high-level software …
Accès ouvert
2025
preprint
OpenAlex
John Jeongseok Yang, K. Lieret, Joyce Yang, Carlos E. Jimenez-Gomez et autres
Current benchmarks for coding evaluate language models (LMs) on concrete, well-specified tasks such as fixing specific bugs or writing targeted tests. However, human programmers do not spend all day incessantly addressing isolated tasks. Instead, real-world software development is grounded in the pursuit …
Accès ouvert
2025
preprint
OpenAlex
Minhui Zhu, Minyang Tian, Xiaocheng Yang, Tianci Zhou et autres
While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason effectively through complex, open-ended challenges found in frontier physics research? And crucially, what kinds of reasoning tasks do physicists want LLMs to …
Accès ouvert
2025
conference-paper
OpenAlex
Ori Press, Brandon D. Amos, Haoyu Zhao, Yikai Wu et autres
Despite progress in language model (LM) capabilities, evaluations have thus far focused on models' performance on tasks that humans have previously solved, including in programming (Jimenez et al., 2024) and mathematics (Glazer et al., 2024). We therefore propose testing models' ability to …
de, us, il, ch
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
John Jeongseok Yang, K. Lieret, Carlos E. Jimenez-Gomez, Alexander Wettig et autres
Despite recent progress in Language Models (LMs) for software engineering, collecting training data remains a significant pain point. Existing datasets are small, with at most 1,000s of training instances from 11 or fewer GitHub repositories. The procedures to curate such datasets are …
us
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
John Jeongseok Yang, Carlos E. Jimenez-Gomez, Alex Zhang, K. Lieret et autres
Autonomous systems for software engineering are now capable of fixing bugs and developing features. These systems are commonly evaluated on SWE-bench (Jimenez et al., 2024a), which assesses their ability to solve software issues from GitHub repositories. However, SWE-bench uses only Python repositories, …
Accès ouvert
2024
preprint
OpenAlex
Talor Abramovich, Meet Udeshi, Minghao Shao, K. Lieret et autres
Although language model (LM) agents have demonstrated increased performance in multiple domains, including coding and web-browsing, their success in cybersecurity has been limited. We present EnIGMA, an LM agent for autonomously solving Capture The Flag (CTF) challenges. We introduce new tools and …
Accès ouvert
2024
preprint
OpenAlex
Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin et autres
Language agents, built on top of language models (LMs), are systems that can interact with complex environments, such as the open web. In this work, we examine whether such agents can perform realistic and time-consuming tasks on the web, e.g., monitoring real-estate …
Accès ouvert
2024
preprint
OpenAlex
Minyang Tian, S Zhang, Xinan Chen, C. Fan et autres
Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this issue by examining LMs' capabilities to generate code for solving real scientific research problems. Incorporating …
Accès ouvert
2024
preprint
OpenAlex
Ori Press, Andreas Hochlehnert, Ameya Prabhu, Vishaal Udandarao et autres
Thousands of new scientific papers are published each month. Such information overload complicates researcher efforts to stay current with the state-of-the-art as well as to verify and correctly attribute claims. We pose the following research question: Given a text excerpt referencing a …
Accès ouvert
2024
preprint
OpenAlex
John Jeongseok Yang, Carlos E. Jimenez-Gomez, Alexander Wettig, K. Lieret et autres
Language model (LM) agents are increasingly being used to automate complicated tasks in digital environments. Just as humans benefit from powerful software applications, such as integrated development environments, for complex tasks like software engineering, we posit that LM agents represent a new …
Accès ouvert
2024
conference-paper
OpenAlex
Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin et autres
Language agents, built on top of language models (LMs), are systems that can interact with complex environments, such as the open web.In this work, we examine whether such agents can perform realistic and time-consuming tasks on the web, e.g., monitoring real-estate markets …
il, us
(code pays fourni par la source)