Accès ouvert
2026
preprint
OpenAlex
Vinay Samuel, Varun Ursekar, Vijay S. Kalmath, Apaar Shanker et autres
Before an LLM agent tackles tasks in a new environment, it can inspect available corpora and tools and construct reusable resources such as indices, scripts, or procedural guidance. Most automated adaptation methods, however, rely on task examples, trajectories, or evaluation feedback to …
Accès ouvert
2026
preprint
OpenAlex
Veronica Chatrath, Bryan Zhu, Jingxuan Fan, George Pu et autres
An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a required reliability level, …
in
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Veronica Chatrath, Bryan Zhu, George Pu, Jingxuan Fan et autres
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring …
Accès ouvert
2026
preprint
OpenAlex
Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser et autres
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and …
Accès ouvert
2026
preprint
OpenAlex
Akshay Manglik, Apaar Shanker, Kaustubh Deshpande, Jason Qin et autres
Diagnosing failures in LLM agents remains largely manual. Practitioners inspect a small subset of execution traces, form ad-hoc hypotheses, and iterate. This process misses patterns that only emerge across trace populations and does not scale to production corpora where individual traces span …
Accès ouvert
2026
preprint
OpenAlex
Varun Ursekar, Apaar Shanker, Veronica Chatrath, Yuan et autres
An important emerging application of coding agents is agent harness optimization: the iterative improvement of a target agent by editing and evaluating its code. Despite its relevance, the community lacks a systematic understanding of coding agent performance on this task. Harness optimization …
Accès ouvert
2025
preprint
OpenAlex
Julia Shuieh, Prasann Singhal, Apaar Shanker, J. Heyer et autres
Supervised and preference-based fine-tuning techniques have become popular for aligning large language models (LLMs) with user intent and correctness criteria. However, real-world training data often exhibits spurious correlations -- arising from biases, dataset artifacts, or other "shortcut" features -- that can compromise …
Accès ouvert
2024
preprint
OpenAlex
Yung-Chieh Chan, George Pu, Apaar Shanker, Parth Suresh et autres
As large language models (LLMs) are applied to more use cases, creating high quality, task-specific datasets for fine-tuning becomes a bottleneck for model improvement. Using high quality human data has been the most common approach to unlock model performance, but is prohibitively …