RRSECS: Referring remote sensing expression comprehension and segmentation
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Understanding and interpreting a specific object from large-scale remote sensing (RS) scenes provide basic support for various practical applications. To achieve it, visual grounding (VG) and referring image segmentation (RIS) are two main techniques that aim to localize and segment the referred object given a free-form linguistic expression. Currently, most works of VG and RIS focus on natural images, with only a few generalizing to RS images and further developing remote sensing visual grounding (RSVG) and referring remote sensing image segmentation (RRSIS). However, RSVG and RRSIS are designed to solve separate tasks, ignoring the benefits of jointly learning localization and segmentation. In this work, we introduce the task of referring remote sensing expression comprehension and segmentation (RRSECS) to explore the potential of multi-task learning in vision-language understanding. Specifically, we construct the first benchmark for this task, namely RefDIOR, which contains image-expression-box-mask quadruplets for training and evaluating different models, enabling us to advance the research of RRSECS. Then, we benchmark extensive methods across VG, RSVG, RIS, and RRSIS on RefDIOR, and give insightful analyses of their performances and limitations. Finally, we propose a novel cross-task collaborative Transformer (CCFormer) to accomplish language-guided end-to-end localization and segmentation, serving as a strong baseline for RRSECS. CCFormer consists of the multi-scale cross-modal fusion module to obtain fine-grained aligned vision-language features, the language-aware gated decoupling module to assign discriminative multi-modal features for each task, and the cross-task collaborative loss to refine multiple outputs in a co-rectify manner. Experimental results on RefDIOR demonstrate the effectiveness and generalization of CCFormer in addressing the challenges of RSVG and RRSIS. The code and dataset will be publicly accessible athttps://github.com/IPIU-XDU/RSFM.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- RRSECS: Referring remote sensing expression comprehension and segmentation
- Date Crossref
- 01/09/2025
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.