Aller au contenu principal
Accès ouvert déclaré 2025 conference-paper

Unspoken Details: Inferring Hidden Causality and Retrieving Domain-Specific Knowledge for Image Generation

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Text-to-image (T2I) generation has advanced significantly in recent years, yet current models often struggle with prompts that imply causal sequences or require knowledge of culturally grounded entities. This limitation stems from a fundamental "semantic gap" between a user’s rich intent and the model’s statistical interpretation of text. To address these limitations, we propose a causality-aware multimodal framework that integrates large language models (LLMs), visual-language verification, and domain-specific image retrieval within an iterative, self-correcting pipeline. The system first decomposes prompts into structured representations of causal chains and named entities. It then retrieves aligned visual references from a multimodal knowledge base to ground these abstract concepts. These components are fused into an enriched multimodal prompt for a frozen-backbone diffusion model. A verification module, powered by a Vision-Language Model (VLM), evaluates the causal and semantic consistency of the generated output, triggering a refinement loop when necessary. This closed-loop design enables more coherent, grounded, and context-sensitive image synthesis, particularly in complex or culturally nuanced scenarios. Our approach expands the expressive capacity of T2I systems by explicitly modeling and integrating the unspoken details of physical logic and domain knowledge, thereby bridging the semantic gap and producing images that are more faithful to user intent.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Unspoken Details: Inferring Hidden Causality and Retrieving Domain-Specific Knowledge for Image Generation
Date Crossref
01/12/2025
Éditeur
ACM
Type
proceedings-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • The Hong Kong University of Science and Technology (Guangzhou) pays non établi dans la notice
    Université ou école supérieure

The Hong Kong University of Science and Technology (Guangzhou).

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Multimodal Machine Learning ApplicationsGenerative Adversarial Networks and Image SynthesisTopic Modeling

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.