Accès ouvert
2026
preprint
OpenAlex
Utkarsh Tyagi, R. K. Selvakumar, Advait Gosai, Sonal Kumar et autres
Full-duplex spoken dialogue models support low-latency turn taking, interruption handling, and backchanneling, yet a key capability remains underexplored: steerability, the ability to reliably shift conversational behavior along attributes such as tone, persona, speaking rate, and voice style in response to user instructions. …
Accès ouvert
2026
preprint
OpenAlex
Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang et autres
The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the …
Accès ouvert
2026
preprint
OpenAlex
Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang et autres
The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the …
Accès ouvert
2026
preprint
OpenAlex
Veronica Chatrath, Bryan Zhu, Jingxuan Fan, George Pu et autres
An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a required reliability level, …
in
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang et autres
Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer. The rubric, however, is a fixed proxy for quality, never a complete description of it, …
Accès ouvert
2026
preprint
OpenAlex
Harsh Raj, Vipul Gupta, Anas Mahmoud, Razvan-Gabriel Dumitru et autres
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engineering, environment redesign, or benchmark …
Accès ouvert
2026
preprint
OpenAlex
Harsh Raj, Vipul Gupta, Anas Mahmoud, Razvan-Gabriel Dumitru et autres
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engineering, environment redesign, or benchmark …
Accès ouvert
2026
preprint
OpenAlex
MohammadHossein Rezaei, Anas Mahmoud, Zihao Wang, Utkarsh Tyagi et autres
Rubrics have emerged as an alternative to RLVR in open-ended domains where a single ground-truth final answer is not available. Existing rubric-based training methods rely on an LLM verifier that scores each rollout against rubrics. This introduces substantial training-time overhead, exposes optimization …
Accès ouvert
2026
preprint
OpenAlex
Utkarsh Tyagi, Xingang Guo, MohammadHossein Rezaei, Daniel George et autres
Reinforcement learning with verifiable rewards has made post-training highly effective when correctness can be checked automatically. However, many important model behaviors require satisfying several qualitative criteria at once. Rubric-based rewards address this setting by grading prompt-specific criteria and aggregating them into a …
Accès ouvert
2026
preprint
OpenAlex
Utkarsh Tyagi, Xingang Guo, MohammadHossein Rezaei, Daniel George et autres
Reinforcement learning with verifiable rewards has made post-training highly effective when correctness can be checked automatically. However, many important model behaviors require satisfying several qualitative criteria at once. Rubric-based rewards address this setting by grading prompt-specific criteria and aggregating them into a …
Accès ouvert
2026
preprint
OpenAlex
Anas Mahmoud, MohammadHossein Rezaei, Zihao Wang, Anisha Gunjal et autres
Reinforcement learning with verifiable rewards has enabled strong post-training gains in domains such as math and coding, though many open-ended settings rely on rubric-based rewards. We study reward hacking in rubric-based RL, where a policy is optimized against a training verifier but …
Accès ouvert
2026
preprint
OpenAlex
Anas Mahmoud, MohammadHossein Rezaei, Zihao Wang, Anisha Gunjal et autres
Reinforcement learning with verifiable rewards has enabled strong post-training gains in domains such as math and coding, though many open-ended settings rely on rubric-based rewards. We study reward hacking in rubric-based RL, where a policy is optimized against a training verifier but …