Enhanced 3D Object Detection for Smart Cities with Multimodal Approaches for Autonomous Driving in Dynamic Environme
Le résumé fourni par la source
With the continuous development of Intelligent Transportation Systems (ITS), autonomous driving technologies have become increasingly prevalent in modern transportation. However, existing 3D object detection methods still suffer from depth information loss and misalignment between different modalities, leading to insufficient detection accuracy and robustness. To address these challenges, this paper proposes a novel multimodal 3D object detection model, termed FusionMultiNet. Specifically, FusionMultiNet integrates geometric information from LiDAR with semantic features from RGB images. By introducing the Multi-Depth Unprojection (MDU) strategy and the Gated Modality-Aware Convolution (GMA-Conv) module, the proposed model effectively resolves the issue of depth information loss while jointly alleviating feature fusion biases caused by modality misalignment. The MDU strategy achieves modality alignment in physical space through multi-depth unprojection, whereas the GMA-Conv module adaptively fuses image semantic features under the guidance of LiDAR geometric cues, significantly improving detection accuracy and robustness. In addition, FusionMultiNet incorporates a Modality-Specific Context Encoder (MSCE) to further enhance object feature representation. Experimental results demonstrate that FusionMultiNet consistently outperforms existing state-of-the-art methods on the nuScenes and KITTI benchmark datasets. The proposed model not only improves detection accuracy but also substantially enhances robustness in complex and dynamic traffic scenarios, providing a more reliable solution for autonomous driving in intelligent transportation systems.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.