DKARNet: Prompt-Guided Knowledge-Driven Multimodal Reasoning Network for UAV Object Detection
Résumé fourni par la source
While multimodal reasoning models hold potential for advancing remote sensing toward unified, language-centric frameworks, their application to unmanned aerial vehicle (UAV) imagery faces severe challenges, primarily extreme scale variations and complex background distractors. Recent studies attempt to bridge this gap using vision foundation models and vision–language priors to provide broader semantic context. However, directly aligning these global representations with local CNN features can trigger semantic dilution, where patch-based self-attention smooths out the high-frequency local textures of microinstances. Furthermore, current relational reasoning based solely on cross-modal semantic similarity ignores large spatial distances, frequently generating false positive connections with distant background noise. To establish a robust perceptual foundation for multimodal reasoning in aerial scenes, we propose a prompt-guided dynamic knowledge-driven multimodal adaptive reasoning network (DKARNet). The framework comprises a semantic grounding module (SGM) and a dynamic adaptive reasoning module (DARM). The SGM introduces a multiscale spatial attention mechanism that integrates DINOv3’s global priors while shielding the spatially precise CNN features from attention-induced smoothing. Building on this grounding, the DARM functions as a multimodal reasoning engine to construct a joint spatial–semantic knowledge graph. Under dynamic prompt guidance, it performs adaptive reasoning by applying a Gaussian spatial penalty to explicitly remove false connections between foreground microtargets and far-field distractors. Extensive experiments demonstrate that DKARNet achieves state-of-the-art performance, outperforming existing methods by 3.4% and 6.2% on mAP50-95 and mAP75, respectively, on the VisDrone 2019 dataset, and showing substantial improvements on the UAVDT dataset. This prompt-guided spatial–semantic alignment provides a critical multimodal grounding step, enabling future reasoning language models to execute complex agentic decisions in UAV environments.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- DKARNet: Prompt-Guided Knowledge-Driven Multimodal Reasoning Network for UAV Object Detection
- Date Crossref
- 01/01/2026
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.