SATA-Bench : Select All That Apply Benchmark For Multiple Choice Questions
Weijie Xu, shixian cui, Xi Fang, Chi Xue et autres
us (code pays fourni par la source)
Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.
Weijie Xu, shixian cui, Xi Fang, Chi Xue et autres
us (code pays fourni par la source)
Xi Fang, Weijie Xu, Yingqiang Ge, Yuhui Xu et autres
Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future …
Xi Fang, Weijie Xu, Yingqiang Ge, Yuhui Xu et autres
Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future …
Association for Computational Linguistics 2026, Stephanie Eckman, Xi Fang, Chandan K. Reddy et autres
When an AI assistant remembers that Sarah is a single mother working two jobs, does it interpret her stress differently than if she were a wealthy executive? As personalized AI systems increasingly incorporate long-term user memory, understanding how this memory shapes emotional …
de, us (code pays fourni par la source)
Robert Chew, Stephanie Eckman, Christoph Kern, Frauke Kreuter
Supervised machine learning assumes that labeled data provide accurate measurements of the concepts models are meant to learn. Yet in practice, human labeling introduces systematic variation arising from ambiguous items, divergent interpretations, and simple mistakes. Machine learning research commonly treats all disagreement …
Robert Chew, Stephanie Eckman, Christoph Kern, Frauke Kreuter
Supervised machine learning assumes that labeled data provide accurate measurements of the concepts models are meant to learn. Yet in practice, human labeling introduces systematic variation arising from ambiguous items, divergent interpretations, and simple mistakes. Machine learning research commonly treats all disagreement …
us, de (code pays fourni par la source)
Jacob Beck, Stephanie Eckman, Christoph Kern, Frauke Kreuter
Human-AI collaboration increasingly drives decision-making across industries. While AI systems promise efficiency gains by providing automated suggestions for human review, these workflows can trigger cognitive biases that degrade performance. This paper reveals the psychological factors that determine when these collaborations succeed or …
de, us (code pays fourni par la source)
Xi Fang, Weijie Xu, Yuchong Zhang, Scott Nickleach et autres
Xi Fang, Weijie Xu, Yuchong Zhang, Scott Nickleach, Stephanie Eckman, Chandan K. Reddy. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2026.
Elissa Scherer, Kelly R. Evenson, Carmen C. Cuthbertson, Stephanie Eckman et autres
us (code pays fourni par la source)
Akoi Zoumanigui, Beido Nassirou, Rob Henry, Scott D. Nash et autres
Population-based prevalence surveys are essential for decision-making on interventions to achieve trachoma elimination as a public health problem. This paper outlines the methodologies of Tropical Data, which supports work to undertake those surveys.Tropical Data is a consortium of partners that supports health …
us (code pays fourni par la source)
Jacob Beck, Stephanie Eckman, Christoph Kern, Frauke Kreuter
Human-AI collaboration increasingly drives decision-making across industries, from medical diagnosis to content moderation. While AI systems promise efficiency gains by providing automated suggestions for human review, these workflows can trigger cognitive biases that degrade performance. We know little about the psychological factors …
Weijie Xu, Xi Fang, Xue Chi, Stephanie Eckman et autres
Large language models (LLMs) are increasingly evaluated on single-answer multiple-choice tasks, yet many real-world problems require identifying all correct answers from a set of options. This capability remains underexplored. We introduce SATA-BENCH, the first dedicated benchmark for evaluating LLMs on Select All …
BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.
L'essentiel de l'actu tech du Burkina & d'Afrique, chaque semaine dans votre boîte mail.
Gratuit · sans spam · désinscription en un clic