Accès ouvert
2026
preprint
OpenAlex
Bofeng Huang, Jacques Sun, Diane Bouchacourt, Nicolas Barascud et autres
Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining. Encoders for dense-terminology domains such as medicine, by contrast, are pretrained on small, manually-curated corpora that limit scalability and writing style diversity, a bottleneck even more severe in …
Accès ouvert
2026
preprint
OpenAlex
Bofeng Huang, Jacques Sun, Diane Bouchacourt, Nicolas Barascud et autres
Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining. Encoders for dense-terminology domains such as medicine, by contrast, are pretrained on small, manually-curated corpora that limit scalability and writing style diversity, a bottleneck even more severe in …
Accès ouvert
2024
preprint
OpenAlex
Samuel J. Bell, Diane Bouchacourt, Levent Sagun
Neural networks can fail when the data contains spurious correlations. To understand this phenomenon, researchers have proposed numerous spurious correlations benchmarks upon which to evaluate mitigation methods. However, we observe that these benchmarks exhibit substantial disagreement, with the best methods on one …
Accès ouvert
2024
preprint
OpenAlex
Haider Al-Tahan, Quentin Garrido, Randall Balestriero, Diane Bouchacourt et autres
Significant research efforts have been made to scale and improve vision-language model (VLM) training approaches. Yet, with an ever-growing number of benchmarks, researchers are tasked with the heavy burden of implementing each protocol, bearing a non-trivial computational cost, and making sense of …
Accès ouvert
2024
preprint
OpenAlex
Vlad Sobal, Mark Ibrahim, Randall Balestriero, Vivien Cabannes et autres
Learning good representations involves capturing the diverse ways in which data samples relate. Contrastive loss - an objective matching related samples - underlies methods from self-supervised to multimodal learning. Contrastive losses, however, can be viewed more broadly as modifying a similarity graph …
Accès ouvert
2024
preprint
OpenAlex
O. Kitouni, N. S. Nolte, Diane Bouchacourt, Adina Williams et autres
Today's best language models still struggle with hallucinations: factually incorrect generations, which impede their ability to reliably retrieve information seen during training. The reversal curse, where models cannot recall information when probed in a different order than was encountered during training, exemplifies …
Accès ouvert
2024
conference-paper
OpenAlex
Mazda Moayeri, Michael Rabbat, Mark Ibrahim, Diane Bouchacourt
Vision-language models enable open-world classification of objects without the need for any retraining. While this zero-shot paradigm marks a significant advance, even today’s best models exhibit skewed performance when objects are dissimilar from their typical depiction. Real world objects such as pears …
Accès ouvert
2024
preprint
OpenAlex
Florian Bordes, Richard Yuanzhe Pang, Anurag Ajay, Alexander C. Li et autres
Following the recent popularity of Large Language Models (LLMs), several attempts have been made to extend them to the visual domain. From having a visual assistant that could guide us through unfamiliar environments to generative models that produce images using only a …
2024
conference-paper
OpenAlex
O. Kitouni, N. S. Nolte, Diane Bouchacourt, Adina Williams et autres
2024
conference-paper
OpenAlex
Haider Al-Tahan, Quentin Garrido, Randall Balestriero, Diane Bouchacourt et autres
Accès ouvert
2023
preprint
OpenAlex
Polina Kirichenko, Mark Ibrahim, Randall Balestriero, Diane Bouchacourt et autres
Data augmentation (DA) encodes invariance and provides implicit regularization critical to a model's performance in image classification tasks. However, while DA improves average accuracy, recent studies have shown that its impact can be highly class dependent: achieving optimal average accuracy comes at …
Accès ouvert
2023
preprint
OpenAlex
Cian Eastwood, Julius von Kügelgen, Linus Ericsson, Diane Bouchacourt et autres
Self-supervised representation learning often uses data augmentations to induce some invariance to "style" attributes of the data. However, with downstream tasks generally unknown at training time, it is difficult to deduce a priori which attributes of the data are indeed "style" and …