Diffusion-Augmented Adaptive Image–Text Contrastive Learning for unsupervised multimodal image clustering
Rattachement africain : cn, au. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Unsupervised image clustering remains challenging because visually similar categories often lack explicit semantic boundaries, while purely visual representations are easily affected by appearance variations, background noise, and ambiguous local structures. Although recent image–text clustering methods introduce external textual semantics through vision–language models and lexical knowledge bases, their contrastive optimization still largely depends on static visual views and manually designed augmentations, which may provide insufficient or noisy positive samples. To address these limitations, we propose Diffusion-Augmented Adaptive Image–Text Contrastive Learning (DAITC), a unified unsupervised multimodal clustering framework that integrates adaptive text counterpart construction with diffusion-based semantic positive generation. Specifically, CLIP is first used to extract image and text embeddings, and a compact semantic vocabulary is constructed from WordNet according to image-level semantic centers. An adaptive temperature mechanism is then introduced to generate image-specific text counterparts by dynamically weighting candidate noun embeddings according to their similarity distributions. Beyond conventional visual augmentation, we further design a text-guided conditional diffusion module that generates semantically consistent positive views conditioned on both visual embeddings and adaptive text counterparts. These diffusion-augmented samples are incorporated into a reliability-aware soft contrastive objective, together with direct image–text alignment and neighborhood-level cross-modal consistency constraints. Finally, cluster balance and confidence regularization are employed to obtain stable and discriminative cluster assignments. Experiments are conducted on STL-10, CIFAR-10, CIFAR-20, DTD, and UCF-101 to evaluate clustering accuracy, semantic alignment, robustness, and generalization ability. The proposed framework provides a reliable way to combine external textual semantics and generative positive augmentation for unsupervised multimodal image clustering.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Diffusion-Augmented Adaptive Image–Text Contrastive Learning for unsupervised multimodal image clustering
- Date Crossref
- 01/11/2026
- Éditeur
- Elsevier BV
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.