Aller au contenu principal
Profil bibliographique

Diane Bouchacourt

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

50Publications signalées
764Citations signalées
0Affiliations récentes

Les domaines associés

Topic ModelingNatural Language Processing TechniquesDomain Adaptation and Few-Shot LearningLanguage and cultural evolutionGenerative Adversarial Networks and Image Synthesis

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

Bofeng Huang, Jacques Sun, Diane Bouchacourt, Nicolas Barascud et autres

Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining. Encoders for dense-terminology domains such as medicine, by contrast, are pretrained on small, manually-curated corpora that limit scalability and writing style diversity, a bottleneck even more severe in …

0 citations HAL (Le Centre pour la Communication Scientifique Directe)
Accès ouvert 2026 preprint OpenAlex

Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

Bofeng Huang, Jacques Sun, Diane Bouchacourt, Nicolas Barascud et autres

Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining. Encoders for dense-terminology domains such as medicine, by contrast, are pretrained on small, manually-curated corpora that limit scalability and writing style diversity, a bottleneck even more severe in …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

Reassessing the Validity of Spurious Correlations Benchmarks

Samuel J. Bell, Diane Bouchacourt, Levent Sagun

Neural networks can fail when the data contains spurious correlations. To understand this phenomenon, researchers have proposed numerous spurious correlations benchmarks upon which to evaluate mitigation methods. However, we observe that these benchmarks exhibit substantial disagreement, with the best methods on one …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Haider Al-Tahan, Quentin Garrido, Randall Balestriero, Diane Bouchacourt et autres

Significant research efforts have been made to scale and improve vision-language model (VLM) training approaches. Yet, with an ever-growing number of benchmarks, researchers are tasked with the heavy burden of implementing each protocol, bearing a non-trivial computational cost, and making sense of …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

$\mathbb{X}$-Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs

Vlad Sobal, Mark Ibrahim, Randall Balestriero, Vivien Cabannes et autres

Learning good representations involves capturing the diverse ways in which data samples relate. Contrastive loss - an objective matching related samples - underlies methods from self-supervised to multimodal learning. Contrastive losses, however, can be viewed more broadly as modifying a similarity graph …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More

O. Kitouni, N. S. Nolte, Diane Bouchacourt, Adina Williams et autres

Today's best language models still struggle with hallucinations: factually incorrect generations, which impede their ability to reliably retrieve information seen during training. The reversal curse, where models cannot recall information when probed in a different order than was encountered during training, exemplifies …

2 citations arXiv (Cornell University)
Accès ouvert 2024 conference-paper OpenAlex

Embracing Diversity: Interpretable Zero-shot Classification Beyond One Vector Per Class

Mazda Moayeri, Michael Rabbat, Mark Ibrahim, Diane Bouchacourt

Vision-language models enable open-world classification of objects without the need for any retraining. While this zero-shot paradigm marks a significant advance, even today’s best models exhibit skewed performance when objects are dissimilar from their typical depiction. Real world objects such as pears …

2 citations
Accès ouvert 2023 preprint OpenAlex

Understanding the Detrimental Class-level Effects of Data Augmentation

Polina Kirichenko, Mark Ibrahim, Randall Balestriero, Diane Bouchacourt et autres

Data augmentation (DA) encodes invariance and provides implicit regularization critical to a model's performance in image classification tasks. However, while DA improves average accuracy, recent studies have shown that its impact can be highly class dependent: achieving optimal average accuracy comes at …

2 citations arXiv (Cornell University)
Accès ouvert 2023 preprint OpenAlex

Self-Supervised Disentanglement by Leveraging Structure in Data Augmentations

Cian Eastwood, Julius von Kügelgen, Linus Ericsson, Diane Bouchacourt et autres

Self-supervised representation learning often uses data augmentations to induce some invariance to "style" attributes of the data. However, with downstream tasks generally unknown at training time, it is difficult to deduce a priori which attributes of the data are indeed "style" and …

0 citations arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.