Aller au contenu principal
Profil bibliographique

Yunzhong He

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

27Publications signalées
139Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

Topic ModelingExplainable Artificial Intelligence (XAI)Reinforcement Learning in RoboticsEthics and Social Impacts of AIArtificial Intelligence in Healthcare and Education

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

SteerDuplex: Steerable Duplex Speech Dialogue Models

Utkarsh Tyagi, R. K. Selvakumar, Advait Gosai, Sonal Kumar et autres

Full-duplex spoken dialogue models support low-latency turn taking, interruption handling, and backchanneling, yet a key capability remains underexplored: steerability, the ability to reliably shift conversational behavior along attributes such as tone, persona, speaking rate, and voice style in response to user instructions. …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang et autres

The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang et autres

The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

READY or Not: Reliable Enterprise Agent Deployment

Veronica Chatrath, Bryan Zhu, Jingxuan Fan, George Pu et autres

An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a required reliability level, …

in (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

Harsh Raj, Vipul Gupta, Anas Mahmoud, Razvan-Gabriel Dumitru et autres

Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engineering, environment redesign, or benchmark …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

Harsh Raj, Vipul Gupta, Anas Mahmoud, Razvan-Gabriel Dumitru et autres

Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engineering, environment redesign, or benchmark …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers

MohammadHossein Rezaei, Anas Mahmoud, Zihao Wang, Utkarsh Tyagi et autres

Rubrics have emerged as an alternative to RLVR in open-ended domains where a single ground-truth final answer is not available. Existing rubric-based training methods rely on an LLM verifier that scores each rollout against rubrics. This introduces substantial training-time overhead, exposes optimization …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR

Utkarsh Tyagi, Xingang Guo, MohammadHossein Rezaei, Daniel George et autres

Reinforcement learning with verifiable rewards has made post-training highly effective when correctness can be checked automatically. However, many important model behaviors require satisfying several qualitative criteria at once. Rubric-based rewards address this setting by grading prompt-specific criteria and aggregating them into a …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR

Utkarsh Tyagi, Xingang Guo, MohammadHossein Rezaei, Daniel George et autres

Reinforcement learning with verifiable rewards has made post-training highly effective when correctness can be checked automatically. However, many important model behaviors require satisfying several qualitative criteria at once. Rubric-based rewards address this setting by grading prompt-specific criteria and aggregating them into a …

0 citations arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.