Accès ouvert
2026
preprint
OpenAlex
Tian Qin, Junzhe Chen, Yuqing Shi, Tianshu Zhang et autres
Large vision-language models (LVLMs) often hallucinate when language priors dominate weak or ambiguous visual evidence. Existing contrastive decoding methods mitigate this problem by comparing predictions from the original image with those from externally perturbed visual inputs, but such references can introduce off-manifold …
Accès ouvert
2026
preprint
OpenAlex
Tian Qin, Junzhe Chen, Yuqing Shi, Tianshu Zhang et autres
Large vision-language models (LVLMs) often hallucinate when language priors dominate weak or ambiguous visual evidence. Existing contrastive decoding methods mitigate this problem by comparing predictions from the original image with those from externally perturbed visual inputs, but such references can introduce off-manifold …
au, cn, us
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Junzhe Chen, Tianshu Zhang, Shiyu Huang, Yuwei Niu et autres
Recently, Omni-modal large language models (OLLMs) have sparked a new wave of research, achieving impressive results in tasks such as audio-video understanding and real-time environment perception. However, hallucination issues still persist. Similar to the bimodal setting, the priors from the text modality …
2025
conference-paper
OpenAlex
J. Chen, Tianshu Zhang, Shiyu Huang, Y. Niu et autres
Despite the recent breakthroughs achieved by Large Vision Language Models (LVLMs) in understanding and responding to complex visual-textual contexts, their inherent hallucination tendencies limit their practical application in real-world scenarios that demand high levels of precision. Existing methods typically either fine-tune the …
cn
(code pays fourni par la source)
2025
article
OpenAlex
Tianshu Zhang, Kun Qian, Siddhartha Sahai, Yuan Tian et autres
Neural text-to-SQL models, which translate natural language questions (NLQs) into SQL queries given a database schema, have achieved remarkable performance. However, database schemas frequently evolve to meet new requirements. Such schema evolution often leads to performance degradation for models trained on static …
us
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
J. Chen, Tianshu Zhang, Shiyu Huang, Y. Niu et autres
Despite the recent breakthroughs achieved by Large Vision Language Models (LVLMs) in understanding and responding to complex visual-textual contexts, their inherent hallucination tendencies limit their practical application in real-world scenarios that demand high levels of precision. Existing methods typically either fine-tune the …
Accès ouvert
2024
article
OpenAlex
Tianshu Zhang, Yuanyuan Jiang, Mulin Liu, Yingying Jiang et autres
We are currently in the digital era of the 21st century, and the rapid advancement of artificial intelligence has brought new energy to the evolution of museums. Museums must inevitably advance towards digitisation. Therefore, the variety of applications for artificial intelligence and …
cn, ru
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Fan Liu, Tianshu Zhang, Wenwen Dai, Wenwen Cai et autres
Multi-modal (vision-language) models, such as CLIP, are replacing traditional supervised pre-training models (e.g., ImageNet-based pre-training) as the new generation of visual foundation models. These models with robust and aligned semantic representations learned from billions of internet image-text pairs and can be applied …
Accès ouvert
2023
preprint
OpenAlex
Tianshu Zhang, Xiang Yue, Yifei Li, Huan Sun
Semi-structured tables are ubiquitous. There has been a variety of tasks that aim to automatically interpret, augment, and query tables. Current methods often require pretraining on tables or special model architecture design, are restricted to specific table types, or have simplifying assumptions …
Accès ouvert
2023
preprint
OpenAlex
Fan Liu, Tianshu Zhang, Wenwen Dai, Wenwen Cai et autres
cn, hk
(code pays fourni par la source)
Accès ouvert
2023
preprint
OpenAlex
Lingbo Mo, Shijie Chen, Ziru Chen, Xiang Bo Deng et autres
We introduce TacoBot, a user-centered task-oriented digital assistant designed to guide users through complex real-world tasks with multiple steps. Covering a wide range of cooking and how-to tasks, we aim to deliver a collaborative and engaging dialogue experience. Equipped with language understanding, …
Accès ouvert
2023
preprint
OpenAlex
Tianshu Zhang, Changchang Liu, Wei‐Han Lee, Yu Su et autres
This paper studies a new task of federated learning (FL) for semantic parsing, where multiple clients collaboratively train one global model without sharing their semantic parsing data. By leveraging data from multiple clients, the FL paradigm can be especially beneficial for clients …