Accès ouvert
2026
conference-paper
OpenAlex
Zonglin Yang, Xie Tong, Jinjie Ni, Ben Gao et autres
Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses remains unexamined due to the lack of a dedicated benchmark.To address this gap, we introduce the first large-scale benchmark for evaluating LLMs on …
cn, sg, au
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Zijian Wu, Jinjie Ni, Xiangyan Liu, Z. Liu et autres
Vision-language models (VLMs) trained via reinforcement learning with verifiable reward (RLVR) have shown notable progress in scaling test-time compute effectively.In this work, we investigate how synthesized RL data can further improve RLVR.To this end, we propose Syn-thRL-a scalable and guaranteed pipeline for …
sg, hk, es
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Jinjie Ni, Qian Liu, Longxu Dou, Chao Du et autres
Under strictly controlled pre-training settings, we observe a Crossover: when unique data is limited, diffusion language models (DLMs) consistently surpass autoregressive (AR) models by training for more epochs. The crossover shifts later with more or higher-quality data, earlier with larger models, and …
Accès ouvert
2025
preprint
OpenAlex
Jinjie Ni, Qian Liu, Chao Du, Longxu Dou et autres
We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results …
Accès ouvert
2025
preprint
OpenAlex
Zijian Wu, Xiangyan Liu, Xinyuan Zhang, Lingjun Chen et autres
MCP standardizes how LLMs interact with external systems, forming the foundation for general agents. However, existing MCP benchmarks remain narrow in scope: they focus on read-heavy tasks or tasks with limited interaction depth, and fail to capture the complexity and realism of …
Accès ouvert
2025
preprint
OpenAlex
Keyu Duan, Yiran Zhao, Zhili Feng, Jinjie Ni et autres
Large Language Models (LLMs) have been observed to process non-human-readable text sequences, such as jailbreak prompts, often viewed as a bug for aligned LLMs. In this work, we present a systematic investigation challenging this perception, demonstrating that unnatural languages - strings that …
Accès ouvert
2025
conference-paper
OpenAlex
Xiangyan Liu, Jinjie Ni, Zijian Wu, Chao Du et autres
Recent advances in reinforcement learning (RL) have strengthened the reasoning capabilities of vision-language models (VLMs). However, enhancing policy exploration to better scale test-time compute remains largely underexplored. In addition, VLMs continue to struggle with imperfect visual perception, which in turn affects the …
sg, cn
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Qi Hui Jia, Siyu Ren, Ziheng Qin, Fuzhao Xue et autres
Datasets nowadays are generally constructed from multiple sources and using different synthetic techniques, making data de-noising and de-duplication crucial before being used for post-training. In this work, we propose to perform instruction tuning by iterative data selection (\ApproachName{}). We measure the quality …
Accès ouvert
2024
preprint
OpenAlex
Jinjie Ni, Yifan Song, Deepanway Ghosal, Bo Li et autres
Perceiving and generating diverse modalities are crucial for AI models to effectively learn from and engage with real-world signals, necessitating reliable evaluations for their development. We identify two major issues in current evaluations: (1) inconsistent standards, shaped by different communities with varying …
2024
article
OpenAlex
Jiaxing Xu, Jinjie Ni, Yiping Ke
sg
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng et autres
Evaluating large language models (LLMs) is challenging. Traditional ground-truth-based benchmarks fail to capture the comprehensiveness and nuance of real-world queries, while LLM-as-judge benchmarks suffer from grading biases and limited query quantity. Both of them may also become contaminated over time. User-facing evaluation, …
Accès ouvert
2024
preprint
OpenAlex
Fuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni et autres
To help the open-source community have a better understanding of Mixture-of-Experts (MoE) based large language models (LLMs), we train and release OpenMoE, a series of fully open-sourced and reproducible decoder-only MoE LLMs, ranging from 650M to 34B parameters and trained on up …