Accès ouvert
2026
preprint
OpenAlex
Yibin Huang, Jixiang Hong, Zongzhao Li, Yuhan Dai et autres
Latents from vision foundation models (VFMs) are semantically rich and well suited for visual understanding. Recent representation autoencoder methods such as RAE have shown that they can provide promising latent spaces for image generation. However, VFM latents remain difficult to model directly: …
Accès ouvert
2026
preprint
OpenAlex
Yibin Huang, Jixiang Hong, Zongzhao Li, Yuhan Dai et autres
Latents from vision foundation models (VFMs) are semantically rich and well suited for visual understanding. Recent representation autoencoder methods such as RAE have shown that they can provide promising latent spaces for image generation. However, VFM latents remain difficult to model directly: …
us, cn
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Mingze Li, Yu Rong, Songyou Li, Lihong Wang et autres
Artificial intelligence has accelerated materials discovery through high-throughput prediction and generation, yet the decision problem remains a formidable bottleneck. While current AI systems readily propose millions of candidates, navigating the decision regarding a viable experimental target requires resolving multi-dimensional judgments across atomic-scale …
Accès ouvert
2026
preprint
OpenAlex
Mingze Li, Yu Rong, Songyou Li, Lihong Wang et autres
Artificial intelligence has accelerated materials discovery through high-throughput prediction and generation, yet the decision problem remains a formidable bottleneck. While current AI systems readily propose millions of candidates, navigating the decision regarding a viable experimental target requires resolving multi-dimensional judgments across atomic-scale …
Accès ouvert
2026
preprint
OpenAlex
Yuelin Zhang, Sijie Cheng, Chen Li, Zongzhao Li et autres
Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models (VLMs) based methods primarily leverage their video understanding capabilities, while neglecting their complex reasoning potential. Furthermore, processing long video trajectories …
Accès ouvert
2026
preprint
OpenAlex
Yuelin Zhang, Sijie Cheng, Chen Li, Zongzhao Li et autres
Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models (VLMs) based methods primarily leverage their video understanding capabilities, while neglecting their complex reasoning potential. Furthermore, processing long video trajectories …
cn
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Zongzhao Li, Xiangzhe Kong, Jiahui Su, Zeyu Ma et autres
cn
(code pays fourni par la source)
Accès ouvert
2025
article
OpenAlex
Jiaqi Han, Jiacheng Cen, Liming Wu, Zongzhao Li et autres
Abstract Geometric graphs are a special kind of graph with geometric features, which are vital to model many scientific problems. Unlike generic graphs, geometric graphs often exhibit physical symmetries of translations, rotations, and reflections, making them ineffectively processed by current Graph Neural …
cn, us
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Zongzhao Li, Zongyang Ma, Mingze Li, Songyou Li et autres
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across diverse tasks, yet they lag significantly behind humans in spatial reasoning. We investigate this gap through Transformation-Driven Visual Reasoning (TVR), a challenging task requiring identification of object transformations across images under varying …
Accès ouvert
2025
preprint
OpenAlex
Zongzhao Li, Jiacheng Cen, Bing Su, Wenbing Huang et autres
Accurately predicting 3D structures and dynamics of physical systems is crucial in scientific applications. Existing approaches that rely on geometric Graph Neural Networks (GNNs) effectively enforce $\mathrm{E}(3)$-equivariance, but they often fall in leveraging extensive broader information. While direct application of Large Language …
Accès ouvert
2024
preprint
OpenAlex
Jiaqi Han, Jiacheng Cen, Liming Wu, Zongzhao Li et autres
Geometric graphs are a special kind of graph with geometric features, which are vital to model many scientific problems. Unlike generic graphs, geometric graphs often exhibit physical symmetries of translations, rotations, and reflections, making them ineffectively processed by current Graph Neural Networks …
Accès ouvert
2023
preprint
OpenAlex
Zongzhao Li, Xiangyu Zhu, Xi Zhang, Zhaoxiang Zhang et autres
How to select relevant key objects and reason about the complex relationships cross vision and linguistic domain are two key issues in many multi-modality applications such as visual question answering (VQA). In this work, we incorporate the visual commonsense information and propose …