Accès ouvert
2025
preprint
OpenAlex
Sina J. Semnani, Jirayu Burapacheep, Arpandeep Khatua, Thanawan Atchariyachanvanit et autres
Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RAG) systems. Ensuring its accuracy is therefore critical. But how accurate is Wikipedia, and how can we …
Accès ouvert
2025
preprint
OpenAlex
Sina J. Semnani, Pingyue Zhang, Wanyue Zhai, Haozhuo Li et autres
This paper presents LEMONADE, a large-scale conflict event dataset comprising 39,786 events across 20 languages and 171 countries, with extensive coverage of region-specific entities. LEMONADE is based on a partially reannotated subset of the Armed Conflict Location & Event Data (ACLED), which …
Accès ouvert
2025
conference-paper
OpenAlex
Sina J. Semnani, Pingyue Zhang, Wanyue Zhai, Haozhuo Li et autres
This paper presents LEMONADE, a large-scale conflict event dataset comprising 39,786 events across 20 languages and 171 countries, with extensive coverage of region-specific entities.LEMONADE is based on a partially reannotated subset of the Armed Conflict Location & Event Data (ACLED), which has …
us
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Sina J. Semnani, Han Zhang, Xiaodong He, Merve Tekgürler et autres
Accès ouvert
2025
conference-paper
OpenAlex
Sina J. Semnani, Jirayu Burapacheep, Arpandeep Khatua, Thanawan Atchariyachanvanit et autres
Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RAG) systems.Ensuring its accuracy is therefore critical.But how accurate is Wikipedia, and how can we improve it?We …
us
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Yucheng Jiang, Yijia Shao, Dekun Ma, Sina J. Semnani et autres
While language model (LM)-powered chatbots and generative search engines excel at answering concrete queries, discovering information in the terrain of unknown unknowns remains challenging for users. To emulate the common educational scenario where children/students learn by listening to and participating in conversations …
Accès ouvert
2024
preprint
OpenAlex
Shicheng Liu, Sina J. Semnani, Harold Triedman, Jialiang Xu et autres
Large Language Models (LLMs) have led to significant improvements in the Knowledge Base Question Answering (KBQA) task. However, datasets used in KBQA studies do not capture the true complexity of KBQA tasks. They either have simple questions, use synthetically generated logical forms, …
Accès ouvert
2024
preprint
OpenAlex
Kazuaki Furumai, Roberto Legaspi, Julio Vizcarra, Yudai Yamazaki et autres
Persuasion plays a pivotal role in a wide range of applications from health intervention to the promotion of social good. Persuasive chatbots employed responsibly for social good can be an enabler of positive individual and social change. Existing methods rely on fine-tuning …
Accès ouvert
2024
preprint
OpenAlex
Heidi C. Zhang, Sina J. Semnani, Farhad Ghassemi, Jialiang Xu et autres
We introduce SPAGHETTI: Semantic Parsing Augmented Generation for Hybrid English information from Text Tables and Infoboxes, a hybrid question-answering (QA) pipeline that utilizes information from heterogeneous knowledge sources, including knowledge base, text, tables, and infoboxes. Our LLM-augmented approach achieves state-of-the-art performance on …
Accès ouvert
2024
preprint
OpenAlex
Andrew Lee, Sina J. Semnani, Galo Castillo-López, Gaël de Chalendar et autres
Creating multilingual task-oriented dialogue (TOD) agents is challenging due to the high cost of training data acquisition. Following the research trend of improving training data efficiency, we show for the first time, that in-context learning is sufficient to tackle multilingual TOD. To …
Accès ouvert
2024
conference-paper
OpenAlex
Shicheng Liu, Jialiang Xu, Wesley Tjangnaka, Sina J. Semnani et autres
While most conversational agents are grounded on either free-text or structured knowledge, many knowledge corpora consist of hybrid sources.This paper presents the first conversational agent that supports the full generality of hybrid data access for large knowledge corpora, through a language we …
us
(code pays fourni par la source)
Accès ouvert
2024
conference-paper
OpenAlex
Heidi Zhang, Sina J. Semnani, Farhad Ghassemi, Jialiang Xu et autres
We introduce SPAGHETTI: Semantic Parsing Augmented Generation for Hybrid English information from Text Tables and Infoboxes, a hybrid question-answering (QA) pipeline that utilizes information from heterogeneous knowledge sources, including knowledge base, text, tables, and infoboxes.Our LLM-augmented approach achieves state-ofthe-art performance on the …
us
(code pays fourni par la source)