Accès ouvert
2026
preprint
OpenAlex
Keyang Zhong, Kuo Wang, Peng Liu, Quanlong Zheng et autres
Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments. While Chain-of-Tool-Thought (CoTT) agent systems enable iterative retrieval and inspection, they suffer from …
Accès ouvert
2026
preprint
OpenAlex
Xiaoming Ren, Ru Zhen, Chao Li, Yang Song et autres
Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive interactions. In this technical report, we introduce X-OmniClaw, a unified mobile agent designed for multimodal understanding and interaction in the Android …
Accès ouvert
2026
preprint
OpenAlex
Xiaoming Ren, Ru Zhen, Chao Li, Yang Song et autres
Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive interactions. In this technical report, we introduce X-OmniClaw, a unified mobile agent designed for multimodal understanding and interaction in the Android …
de
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Kuo Wang, Quanlong Zheng, Junlin Xie, Yanhao Zhang et autres
Video Multimodal Large Language Models~(Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the underlying LLMs, existing Video-MLLMs typically exhibit suboptimal performance on long video scenarios. To understand extended input frames, common solutions span …
cn, nl
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Xiaohui Song, Nan Wang, Yafei Liu, Chao Li et autres
In recent years, while cloud-based MLLMs such as QwenVL, InternVL, GPT-4o, Gemini, and Claude Sonnet have demonstrated outstanding performance with enormous model sizes reaching hundreds of billions of parameters, they significantly surpass the limitations in memory, power consumption, and computing capacity of …
Accès ouvert
2025
preprint
OpenAlex
Peng Liu, Xiaoming Ren, Fengkai Liu, Qingsong Xie et autres
Recent advancements in image-to-video (I2V) generation have shown promising performance in conventional scenarios. However, these methods still encounter significant challenges when dealing with complex scenes that require a deep understanding of nuanced motion and intricate object-action relationships. To address these challenges, we …
Accès ouvert
2025
preprint
OpenAlex
Qilin Wu, Quanlong Zheng, Yanhao Zhang, Junlin Xie et autres
With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating video understanding exhibit significant limitations in coverage, task diversity, and scene adaptability. These shortcomings hinder the accurate assessment of …
2025
conference-paper
OpenAlex
Jinguo Luo, Weihong Ren, Quanlong Zheng, Yanhao Zhang et autres
cn, nl
(code pays fourni par la source)
2022
conference-paper
OpenAlex
Jianhao Gao, Quanlong Zheng, Yandong Guo
Shadow removal is an important yet challenging restoration task. State-of-the-art shadow removal methods usually require paired datasets for training. Existing shadow removal datasets lack large-scale quantity and scene diversity. Hence, models trained on such datasets have poor generalization ability. This paper proposes …
cn
(code pays fourni par la source)
Accès ouvert
2022
erratum
OpenAlex
Xiaotian Qiao, Quanlong Zheng, Ying Shan Cao, Rynson W. H. Lau
cn, hk
(code pays fourni par la source)
2022
article
OpenAlex
Xiaotian Qiao, Quanlong Zheng, Ying Shan Cao, Rynson W. H. Lau
cn, hk
(code pays fourni par la source)
2021
conference-paper
OpenAlex
Quanlong Zheng, Xiaotian Qiao, Ying Shan Cao, Shi Zeng Guo et autres
Single-image reflection removal (SIRR) aims to restore the transmitted image given a single image shot through glass or window. Existing methods rely mainly on information extracted from a single image along with some predefined priors, and fail to give satisfying results on …
hk, ky
(code pays fourni par la source)