Aller au contenu principal
Profil bibliographique

Quanlong Zheng

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

17Publications signalées
261Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

Multimodal Machine Learning ApplicationsGenerative Adversarial Networks and Image SynthesisVideo Surveillance and Tracking MethodsImage Retrieval and Classification TechniquesImage Enhancement Techniques

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

Keyang Zhong, Kuo Wang, Peng Liu, Quanlong Zheng et autres

Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments. While Chain-of-Tool-Thought (CoTT) agent systems enable iterative retrieval and inspection, they suffer from …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

Xiaoming Ren, Ru Zhen, Chao Li, Yang Song et autres

Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive interactions. In this technical report, we introduce X-OmniClaw, a unified mobile agent designed for multimodal understanding and interaction in the Android …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

Xiaoming Ren, Ru Zhen, Chao Li, Yang Song et autres

Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive interactions. In this technical report, we introduce X-OmniClaw, a unified mobile agent designed for multimodal understanding and interaction in the Android …

de (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2025 conference-paper OpenAlex

Free-Moref: Instantly Multiplexing Context Perception Capabilities of Video-Mllms Within Single Inference

Kuo Wang, Quanlong Zheng, Junlin Xie, Yanhao Zhang et autres

Video Multimodal Large Language Models~(Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the underlying LLMs, existing Video-MLLMs typically exhibit suboptimal performance on long video scenarios. To understand extended input frames, common solutions span …

cn, nl (code pays fourni par la source)

0 citations
Accès ouvert 2025 preprint OpenAlex

AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model

Xiaohui Song, Nan Wang, Yafei Liu, Chao Li et autres

In recent years, while cloud-based MLLMs such as QwenVL, InternVL, GPT-4o, Gemini, and Claude Sonnet have demonstrated outstanding performance with enormous model sizes reaching hundreds of billions of parameters, they significantly surpass the limitations in memory, power consumption, and computing capacity of …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM

Peng Liu, Xiaoming Ren, Fengkai Liu, Qingsong Xie et autres

Recent advancements in image-to-video (I2V) generation have shown promising performance in conventional scenarios. However, these methods still encounter significant challenges when dealing with complex scenes that require a deep understanding of nuanced motion and intricate object-action relationships. To address these challenges, we …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding

Qilin Wu, Quanlong Zheng, Yanhao Zhang, Junlin Xie et autres

With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating video understanding exhibit significant limitations in coverage, task diversity, and scene adaptability. These shortcomings hinder the accurate assessment of …

0 citations arXiv (Cornell University)
2022 conference-paper OpenAlex

Towards Real-world Shadow Removal with a Shadow Simulation Method and a Two-stage Framework

Jianhao Gao, Quanlong Zheng, Yandong Guo

Shadow removal is an important yet challenging restoration task. State-of-the-art shadow removal methods usually require paired datasets for training. Existing shadow removal datasets lack large-scale quantity and scene diversity. Hence, models trained on such datasets have poor generalization ability. This paper proposes …

cn (code pays fourni par la source)

10 citations 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.