Accès ouvert
2026
preprint
OpenAlex
Mandana Samiei, Eunice Yiu, Anthony GX-Chen, Dexin Lin et autres
A long-standing finding in the causal learning literature is that adults struggle to identify conjunctive causal rules, where an effect requires the simultaneous presence of multiple causes, while performing better in disjunctive settings. However, most demonstrations of this ``conjunctive handicap'' rely on …
Accès ouvert
2026
preprint
OpenAlex
Mandana Samiei, Eunice Yiu, Anthony GX-Chen, Dexin Lin et autres
A long-standing finding in the causal learning literature is that adults struggle to identify conjunctive causal rules, where an effect requires the simultaneous presence of multiple causes, while performing better in disjunctive settings. However, most demonstrations of this ``conjunctive handicap'' rely on …
us, ca
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Anthony GX-Chen, Ankit Anand, Gheorghe Comanici, Zaheer Abbas et autres
Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern applications such as language model fine-tuning or scientific discovery demand diversity. Existing remedies such as entropy regularization or diversity bonuses often require …
Accès ouvert
2026
preprint
OpenAlex
Anthony GX-Chen, Ankit Anand, Gheorghe Comanici, Dr. Zaheer Abbas et autres
Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern applications such as language model fine-tuning or scientific discovery demand diversity. Existing remedies such as entropy regularization or diversity bonuses often require …
us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
AD Jhaveri, Anthony GX-Chen, Ilia Sucholutsky, Eunsol Choi
Confirmation bias, the tendency to seek evidence that supports rather than challenges one's belief, hinders one's reasoning ability. We examine whether large language models (LLMs) exhibit confirmation bias by adapting the rule-discovery study from human psychology: given a sequence of three numbers …
Accès ouvert
2026
preprint
OpenAlex
AD Jhaveri, Anthony GX-Chen, Ilia Sucholutsky, Eunsol Choi
Confirmation bias, the tendency to seek evidence that supports rather than challenges one's belief, hinders one's reasoning ability. We examine whether large language models (LLMs) exhibit confirmation bias by adapting the rule-discovery study from human psychology: given a sequence of three numbers …
2026
article
OpenAlex
Jeff Guo, Junwu Chen, Anthony GX-Chen, Philippe Schwaller
ch, us
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Anthony GX-Chen, Jatin Prakash, Jeff Guo, Rob Fergus et autres
It is commonly believed that optimizing the reverse KL divergence results in "mode seeking", while optimizing forward KL results in "mass covering", with the latter being preferred if the goal is to sample from multiple diverse modes. We show -- mathematically and …
Accès ouvert
2025
preprint
OpenAlex
Anthony GX-Chen, Dongyan Lin, Mandana Samiei, Doina Precup et autres
Language model (LM) agents are increasingly used as autonomous decision-makers which need to actively gather information to guide their decisions. A crucial cognitive skill for such agents is the efficient exploration and understanding of the causal structure of the world -- key …
Accès ouvert
2024
preprint
OpenAlex
Anthony GX-Chen, Kenneth Marino, Rob Fergus
In the face of difficult exploration problems in reinforcement learning, we study whether giving an agent an object-centric mapping (describing a set of items and their attributes) allow for more efficient learning. We found this problem is best solved hierarchically by modelling …
Accès ouvert
2022
preprint
OpenAlex
Wancong Zhang, Anthony GX-Chen, Vlad Sobal, Yann LeCun et autres
Unsupervised visual representation learning offers the opportunity to leverage large corpora of unlabeled trajectories to form useful visual representations, which can benefit the training of reinforcement learning (RL) algorithms. However, evaluating the fitness of such representations requires training RL algorithms which is …
Accès ouvert
2022
conference-paper
OpenAlex
Anthony GX-Chen, Veronica Chelu, Blake Richards, Joëlle Pineau
Estimating value functions is a core component of reinforcement learning algorithms. Temporal difference (TD) learning algorithms use bootstrapping, i.e. they update the value function toward a learning target using value estimates at subsequent time-steps. Alternatively, the value function can be updated toward …
ca
(code pays fourni par la source)