Unspoken Details: Inferring Hidden Causality and Retrieving Domain-Specific Knowledge for Image Generation
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Text-to-image (T2I) generation has advanced significantly in recent years, yet current models often struggle with prompts that imply causal sequences or require knowledge of culturally grounded entities. This limitation stems from a fundamental "semantic gap" between a user’s rich intent and the model’s statistical interpretation of text. To address these limitations, we propose a causality-aware multimodal framework that integrates large language models (LLMs), visual-language verification, and domain-specific image retrieval within an iterative, self-correcting pipeline. The system first decomposes prompts into structured representations of causal chains and named entities. It then retrieves aligned visual references from a multimodal knowledge base to ground these abstract concepts. These components are fused into an enriched multimodal prompt for a frozen-backbone diffusion model. A verification module, powered by a Vision-Language Model (VLM), evaluates the causal and semantic consistency of the generated output, triggering a refinement loop when necessary. This closed-loop design enables more coherent, grounded, and context-sensitive image synthesis, particularly in complex or culturally nuanced scenarios. Our approach expands the expressive capacity of T2I systems by explicitly modeling and integrating the unspoken details of physical logic and domain knowledge, thereby bridging the semantic gap and producing images that are more faithful to user intent.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Unspoken Details: Inferring Hidden Causality and Retrieving Domain-Specific Knowledge for Image Generation
- Date Crossref
- 01/12/2025
- Éditeur
- ACM
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
The Hong Kong University of Science and Technology (Guangzhou) pays non établi dans la noticeUniversité ou école supérieure
The Hong Kong University of Science and Technology (Guangzhou).
Une affiliation ne permet pas de déduire la nationalité d’un auteur.