Visualization Imaging Model Guided Generalizable Real Rainy Image Deraining
Xin Li, Yuxin Feng, Zhe Huang, Fan Zhou et autres
cn (code pays fourni par la source)
Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.
Xin Li, Yuxin Feng, Zhe Huang, Fan Zhou et autres
cn (code pays fourni par la source)
Ke Du, Yimin Peng, Chao Gao, Fan Zhou et autres
DORAEMON is an open-source PyTorch library that unifies visual object modeling and representation learning across diverse scales. A single YAML-driven workflow covers classification, retrieval and metric learning; more than 1000 pretrained backbones are exposed through a timm-compatible interface, together with modular losses, …
Zhuo Su, Jufeng Li, Fuwei Zhang, Yuxin Feng et autres
Existing learning-based dehazing methods perform well on synthetic data but struggle in real scenarios due to the domain gap, causing residual haze and detail loss. To address this, we propose a Multilevel Subspace Distribution Adapter (MSDA) to progressively reduce the feature distribution …
cn, mo (code pays fourni par la source)
Baoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang et autres
The widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and talks, have led to a dramatic increase in …
cn (code pays fourni par la source)
Juepeng Zheng, Yikun Chen, Xudong Mao, Guanghui Yue et autres
Robust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent literature, these 2D representations fall short of capturing sufficient cues that can be exploited to handle viewpoint variations, and capture …
cn, gb (code pays fourni par la source)
Kuo Wang, Quanlong Zheng, Junlin Xie, Yanhao Zhang et autres
Video Multimodal Large Language Models~(Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the underlying LLMs, existing Video-MLLMs typically exhibit suboptimal performance on long video scenarios. To understand extended input frames, common solutions span …
cn, nl (code pays fourni par la source)
Xu Jin, Zhifang Guo, H. Hu, Yunfei Chu et autres
We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relative to single-modal counterparts. Qwen3-Omni matches the performance of same-sized single-modal models within the Qwen series and excels …
Fan Zhou, Mingchao Wang, Huilin Wang, Bo Zhang
This study presents a methodical approach for optimizing stress isolation grooves in off-axis three-reflector systems through sensitivity analysis of structural parameters. Findings reveal that groove depth and spacing are critical parameters determining isolation effectiveness, accounting for 54% of performance variability according to …
Xinzhu Li, Yang Yi, Yikun Chen, Guanghui Yue et autres
Recent advances in LiDAR technology have enabled privacy-preserving gait recognition using 3D point cloud data. However, existing approaches struggle with the inherent challenges of point cloud processing and understanding such as spatial sparsity, irregular sampling, and complex temporal dynamics. In this paper, …
cn, gb (code pays fourni par la source)
Yi Yang, Xinzhu Li, Yufeng Chen, Guanghui Yue et autres
Stylized human motion synthesis remains a fundamental challenge in computer animation and graphics, with a wide spectrum of applications spanning gaming, film production, virtual reality, and beyond. While recent advances in text-driven motion generation have shown promise, existing approaches face critical limitations …
cn, gb (code pays fourni par la source)
Jiao Li, Shaohan Yin, Jing Hou, Xiaohuan Cao et autres
Manual interpretation of CT images for bone metastasis (BM) detection in primary cancer remains challenging. We present an automated Bone Lesion Detection System (BLDS) developed using CT scans from 2518 patients (9177 BMs; 12,824 non-BM lesions) across five hospitals. The system, developed …
cn (code pays fourni par la source)
Xovee Xu, Yifan Zhang, Fan Zhou, Jingkuan Song
Understanding and predicting the popularity of online User-Generated Content (UGC) is critical for various social and recommendation systems. Existing efforts have focused on extracting predictive features and using pre-trained deep models to learn and fuse multimodal UGC representations. However, the dissemination of …
cn (code pays fourni par la source)
BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.
L'essentiel de l'actu tech du Burkina & d'Afrique, chaque semaine dans votre boîte mail.
Gratuit · sans spam · désinscription en un clic