Learning Feature Inversion for Multi-class Anomaly Detection under General-purpose COCO-AD Benchmark
Jiangning Zhang, Chengjie Wang, Xiangtai Li, Guanzhong Tian et autres
cn, sg (code pays fourni par la source)
Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.
Jiangning Zhang, Chengjie Wang, Xiangtai Li, Guanzhong Tian et autres
cn, sg (code pays fourni par la source)
Jiangning Zhang, Junwei Zhu, Teng Hu, Yabiao Wang et autres
Native 4K (2160$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy …
Yuheng Chen, Teng Hu, Jiangning Zhang, Zhucun Xue et autres
Recent advances in text-to-video diffusion models have enabled the generation of high-quality videos conditioned on textual descriptions. However, most existing text-to-video models rely solely on textual conditions, lacking general fine-grained controllability over video generation. To address this challenge, we propose InstanceV, a …
This paper presents a doctoral research focusing on integrating Retrieval-Augmented Generation (RAG) into video-related multimodal tasks. Existing RAG studies predominantly target text, images, or tabular data, overlooking the unique value of video as a knowledge carrier. We address this gap by: 1) …
cn (code pays fourni par la source)
Jiangning Zhang, Beiwen Tian, Hongrui Zhu, Yufei Jia et autres
Visuomotor imitation learning policies enable robots to efficiently acquire manipulation skills from visual demonstrations. However, as scene complexity and visual distractions increase, policies that perform well in simple settings often experience substantial performance degradation. To address this challenge, we propose ImitDiff, a …
cn (code pays fourni par la source)
Jiangning Zhang, Zhucun Xue, Yabiao Wang, Chengjie Wang et autres
This work focuses on developing parameter-efficient and lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Our goal is to set up the new frontier of the 5 M magnitude lightweight model on various downstream tasks. Inverted Residual Block …
cn, sg (code pays fourni par la source)
Zhucun Xue, Jiangning Zhang, Teng Hu, Haoyang He et autres
The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements for high-quality video generation models. For example, the generation of movie-level Ultra-High …
Zhucun Xue, Jiangning Zhang, Xurong Xie, Yuxuan Cai et autres
Multimodal Large Language Models (MLLMs) perform well in video understanding but degrade on long videos due to fixed-length context and weak long-term dependency modeling. Retrieval-Augmented Generation (RAG) can expand knowledge dynamically, yet existing video RAG schemes adopt fixed retrieval paradigms that ignore …
Teng Hu, Jiangning Zhang, Ran Yi, Jieyu Weng et autres
Employing LLMs for visual generation has recently become a research focus. However, the existing methods primarily transfer the LLM architecture to visual generation but rarely investigate the fundamental differences between language and vision. This oversight may lead to suboptimal utilization of visual …
cn (code pays fourni par la source)
Yue Han, Jiangning Zhang, Junwei Zhu, Runze Hou et autres
Multimodal Language Learning Models (MLLMs) have shown remarkable performance in image understanding, generation, and editing, with recent advancements achieving pixel-level grounding with reasoning. However, these models for common objects struggle with fine-grained face understanding. In this work, we introduce the FacePlayGround-240K dataset, …
cn (code pays fourni par la source)
Yabiao Wang, Shuo Wang, Jiangning Zhang, Ke Fan et autres
Human-human motion generation is essential for understanding humans as social beings. Current methods fall into two main categories: single-person-based methods and separate modeling-based methods. To delve into this field, we abstract the overall generation process into a general framework MetaMotion, which consists …
cn (code pays fourni par la source)
Yinan Chen, Jiangning Zhang, Y.J Bi, Xiaobin Hu et autres
Image inversion is a fundamental task in generative models, aiming to map images back to their latent representations to enable downstream applications such as editing, restoration, and style transfer. This paper provides a comprehensive review of the latest advancements in image inversion …
BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.
L'essentiel de l'actu tech du Burkina & d'Afrique, chaque semaine dans votre boîte mail.
Gratuit · sans spam · désinscription en un clic