Accès ouvert
2026
conference-paper
OpenAlex
Zhiwen Tan, Jiaming Huang, Qintong Wu, Hongxuan Zhang et autres
Large Language Models (LLMs), despite their remarkable capabilities, are prone to generating hallucinated or outdated content due to their static internal knowledge. While Retrieval-Augmented Generation (RAG) integrated with Reinforcement Learning (RL) offers a solution, these methods are fundamentally constrained by a single-query …
cn, fr
(code pays fourni par la source)
Accès ouvert
2026
other
OpenAlex
Association for Artificial Intelligence 2026, Jinjie Gu, Jiaming Huang, Zhiwen Tan et autres
Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, while they remain prone to generating hallucinated or outdated responses due to their static internal knowledge. Recent advancements in Retrieval-Augmented Generation (RAG) methods have aimed to enhance models' search and reasoning …
Accès ouvert
2025
preprint
OpenAlex
Siyuan Lu, Zilu Wang, Hongxuan Zhang, Qintong Wu et autres
Large Language Model (LLM) agents show great promise for complex, multi-turn tool-use tasks, but their development is often hampered by the extreme scarcity of high-quality training data. Supervised fine-tuning (SFT) on synthetic data leads to overfitting, whereas standard reinforcement learning (RL) struggles …
Accès ouvert
2025
preprint
OpenAlex
Jiaming Huang, Qintong Wu, Hongxuan Zhang, Chenyi Zhuang et autres
Large Language Models (LLMs), despite their remarkable capabilities, are prone to generating hallucinated or outdated content due to their static internal knowledge. While Retrieval-Augmented Generation (RAG) integrated with Reinforcement Learning (RL) offers a solution, these methods are fundamentally constrained by a single-query …
2024
conference-paper
OpenAlex
Shaowei Wei, Zhengwei Wu, Xin Li, Qintong Wu et autres
cn
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Shaowei Wei, Zhengwei Wu, Xin Yan Li, Qintong Wu et autres
Sequential recommendation methods play a pivotal role in modern recommendation systems. A key challenge lies in accurately modeling user preferences in the face of data sparsity. To tackle this challenge, recent methods leverage contrastive learning (CL) to derive self-supervision signals by maximizing …
2022
conference-paper
OpenAlex
Hao Qian, Qintong Wu, MingHao Li, Zhengwei Wu et autres
Modeling users' historical behaviors is an essential task in many industrial recommender systems. The user interest representation, in previous works, is obtained through the following paradigm: concrete behaviors are firstly embedded as low-dimensional behavior representations, which are then aggregated conditioning on the …
cn
(code pays fourni par la source)
2022
conference-paper
OpenAlex
Zhaoxin Huan, Gongduo Zhang, Xiaolu Zhang, Jun Zhou et autres
There exists the cold-start problem in the recommendation systems when observed user-item interactions are insufficient. To alleviate this problem, most existing works aim to learn globally shared prior knowledge across all items and be fast adapted to a new item with few …
cn
(code pays fourni par la source)
2022
conference-paper
OpenAlex
Zhigang Huangfu, Gong-Duo Zhang, Zhengwei Wu, Qintong Wu et autres
Conversion rate (CVR) prediction is one of the most essential tasks for digital display advertising. In industrial recommender systems, online learning is particularly favored for its capability to capture the dynamic change of data distribution, which often leads to significantly improvement of …
cn
(code pays fourni par la source)
2022
conference-paper
OpenAlex
Hao Qian, Qintong Wu, Peiyan Zhang, Zhiqiang Zhang et autres
Modern recommendation systems introduce the re-ranking stage to optimize the entire list directly. This paper focuses on the design of re-ranking framework in feed to optimally model the mutual influence between items and further promote user engagement. On mobile devices, users browse …
cn
(code pays fourni par la source)