Accès ouvert
2025
preprint
OpenAlex
M. Sharma, Chaithanya Bandi, Clinton Wang, Ankit Aich et autres
Deep Research (DR) is an emerging agent application that leverages large language models (LLMs) to address open-ended queries. It requires the integration of several capabilities, including multi-step reasoning, cross-document synthesis, and the generation of evidence-backed, long-form answers. Evaluating DR remains challenging because …
Accès ouvert
2025
preprint
OpenAlex
Xiang Bo Deng, Jeff Da, Edwin Pan, Yun He et autres
We introduce SWE-Bench Pro, a substantially more challenging benchmark that builds upon the best practices of SWE-BENCH [25], but is explicitly designed to capture realistic, complex, enterprise-level problems beyond the scope of SWE-BENCH. SWE-BENCH PRO contains 1,865 problems sourced from a diverse …
Accès ouvert
2025
preprint
OpenAlex
Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath et autres
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for complex reasoning tasks with clear correctness signals such as math and coding. However, extending it to real-world reasoning tasks is challenging, as evaluation depends on nuanced, multi-criteria judgments rather than binary correctness. …
Accès ouvert
2025
preprint
OpenAlex
Vaskar Nath, Elaine Lau, Anisha Gunjal, Manasi Sharma et autres
We study the process through which reasoning models trained with reinforcement learning on verifiable rewards (RLVR) can learn to solve new problems. We find that RLVR drives performance in two main ways: (1) by compressing pass@$k$ into pass@1 and (2) via "capability …
Accès ouvert
2025
preprint
OpenAlex
Jeff Da, C. Wang, Xiang Bo Deng, Yuntao Ma et autres
Reinforcement Learning from Verifiable Rewards (RLVR) has been widely adopted as the de facto method for enhancing the reasoning capabilities of large language models and has demonstrated notable success in verifiable domains like math and competitive programming tasks. However, the efficacy of …
Accès ouvert
2025
preprint
OpenAlex
Clinton J. Wang, Cristina Menghini, Johannes Mols, Jack Doughty et autres
As language models master existing reasoning benchmarks, we need new challenges to evaluate their cognitive frontiers. Puzzle-solving events are rich repositories of challenging multimodal problems that test a wide range of advanced reasoning and knowledge capabilities, making them a unique testbed for …
Accès ouvert
2025
preprint
OpenAlex
Vaskar Nath, P. Raja, Sean M. Hendryx
Despite recent advances in AI, the development of systems capable of executing complex, multi-step reasoning tasks involving multiple tools remains a significant challenge. Current benchmarks fall short in capturing the real-world complexity of tool-use reasoning, where verifying the correctness of not only …
Accès ouvert
2024
article
OpenAlex
Pratima Khatri‐Chhetri, Hans‐Erik Andersen, Bruce D. Cook, Sean M. Hendryx et autres
The boreal biome, the largest terrestrial biome on Earth, is increasingly vulnerable to climate change due to warming twice as rapidly as the global average. Climate change has increased the temperature, frequency, severity, and amount of area burned, which is leading to …
us
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Spencer Whitehead, Jacob R. Phillips, Sean M. Hendryx
Multimodal language models can exhibit hallucinations in their outputs, which limits their reliability. The ability to automatically detect these errors is important for mitigating them, but has been less explored and existing efforts do not localize hallucinations, instead framing this as a …
Accès ouvert
2024
preprint
OpenAlex
Will LeVine, Benjamin Pikus, Jacob R. Phillips, Berk Norman et autres
As deep neural networks become adopted in high-stakes domains, it is crucial to identify when inference inputs are Out-of-Distribution (OOD) so that users can be alerted of likely drops in performance and calibration despite high confidence -- ultimately to know when networks' …
Accès ouvert
2023
preprint
OpenAlex
Sean M. Hendryx
Foundation models, specifically Large Language Models (LLMs), have lately gained wide-spread attention and adoption. Reinforcement Learning with Human Feedback (RLHF) involves training a reward model to capture desired behaviors, which is then used to align LLM's. These reward models are additionally used …
Accès ouvert
2023
article
OpenAlex
Pratima Khatri‐Chhetri, Liz van Wagtendonk, Sean M. Hendryx, Van R. Kane
Tree mortality is rapidly increasing as a result of more frequent and extensive droughts and forest fires across the globe. The increasing pace and scale of disturbance and resulting mortality necessitates the study of tree mortality at the scale at which it …
us
(code pays fourni par la source)