Accès ouvert
2025
preprint
OpenAlex
Leire Benito-Del-Valle, Artzai Picón, Daniel Múgica, Javier Romero et autres
Herbicide field trials require accurate identification of plant species and assessment of herbicide-induced damage across diverse environments. While general-purpose vision foundation models have shown promising results in complex visual domains, their performance can be limited in agriculture, where fine-grained distinctions between species …
Accès ouvert
2025
preprint
OpenAlex
John Jeongseok Yang, K. Lieret, Joyce Yang, Carlos E. Jimenez-Gomez et autres
Current benchmarks for coding evaluate language models (LMs) on concrete, well-specified tasks such as fixing specific bugs or writing targeted tests. However, human programmers do not spend all day incessantly addressing isolated tasks. Instead, real-world software development is grounded in the pursuit …
2025
book
OpenAlex
Carlos E. Jimenez-Gomez, A.B Shrinivass, Simson Garfinkel
Accès ouvert
2025
preprint
OpenAlex
Quan Shi, Carlos E. Jimenez-Gomez, Stephen Dong, Brian Seo et autres
As language models achieve increasingly human-like capabilities in conversational text generation, a critical question emerges: to what extent can these systems simulate the characteristics of specific individuals? To evaluate this, we introduce IMPersona, a framework for evaluating LMs at impersonating specific individuals' …
2025
book-chapter
OpenAlex
Carlos E. Jimenez-Gomez, Sandra Elena, Peter Sharp Vargas
Accès ouvert
2025
article
OpenAlex
Artzai Picón, Daniel Múgica, Itziar Eguskiza, Arantza Bereciartua et autres
Herbicide research and development necessitate specific trials to monitor the effects of various herbicide formulations, quantities, and protocols on different plant species and growth stages. These trials are necessary to ensure the safety and efficacy of the developed products. Currently, these tests …
Accès ouvert
2025
preprint
OpenAlex
Leire Benito-Del-Valle, Artzai Picón, Daniel Múgica, Javier Romero et autres
es, us, de
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Quan Shi, Carlos E. Jimenez-Gomez, Shunyu Yao, Nick Haber et autres
Recent advancements in AI reasoning have driven substantial improvements across diverse tasks. A critical open question is whether these improvements also yields better knowledge transfer: the ability of models to communicate reasoning in ways humans can understand, apply, and learn from. To …
cn, us
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
John Jeongseok Yang, K. Lieret, Carlos E. Jimenez-Gomez, Alexander Wettig et autres
Despite recent progress in Language Models (LMs) for software engineering, collecting training data remains a significant pain point. Existing datasets are small, with at most 1,000s of training instances from 11 or fewer GitHub repositories. The procedures to curate such datasets are …
us
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Edwin Requena-Zúñiga, Miriam Palomino-Salcedo, María García-Mendoza, Maribel Dana Figueroa-Romero et autres
Abstract This report details the first detection of the Oropouche virus (OROV) in Culicoides insignis in the Ucayali region (Peruvian Amazon) during an outbreak study in 2024. Captures conducted in peri-urban and rural areas showed that C. insignis accounted for 96.7% of …
pe
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
John Jeongseok Yang, Carlos E. Jimenez-Gomez, Alex Zhang, K. Lieret et autres
Autonomous systems for software engineering are now capable of fixing bugs and developing features. These systems are commonly evaluated on SWE-bench (Jimenez et al., 2024a), which assesses their ability to solve software issues from GitHub repositories. However, SWE-bench uses only Python repositories, …
Accès ouvert
2024
preprint
OpenAlex
Talor Abramovich, Meet Udeshi, Minghao Shao, K. Lieret et autres
Although language model (LM) agents have demonstrated increased performance in multiple domains, including coding and web-browsing, their success in cybersecurity has been limited. We present EnIGMA, an LM agent for autonomously solving Capture The Flag (CTF) challenges. We introduce new tools and …