Aller au contenu principal
Profil bibliographique

Jinjie Ni

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

38Publications signalées
527Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

Topic ModelingNatural Language Processing TechniquesSpeech and dialogue systemsMultimodal Machine Learning ApplicationsSpeech Recognition and Synthesis

Les publications récentes

Accès ouvert 2026 conference-paper OpenAlex

ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

Zonglin Yang, Xie Tong, Jinjie Ni, Ben Gao et autres

Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses remains unexamined due to the lack of a dedicated benchmark.To address this gap, we introduce the first large-scale benchmark for evaluating LLMs on …

cn, sg, au (code pays fourni par la source)

1 citation
Accès ouvert 2026 conference-paper OpenAlex

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis

Zijian Wu, Jinjie Ni, Xiangyan Liu, Z. Liu et autres

Vision-language models (VLMs) trained via reinforcement learning with verifiable reward (RLVR) have shown notable progress in scaling test-time compute effectively.In this work, we investigate how synthesized RL data can further improve RLVR.To this end, we propose Syn-thRL-a scalable and guaranteed pipeline for …

sg, hk, es (code pays fourni par la source)

0 citations
Accès ouvert 2025 preprint OpenAlex

Diffusion Language Models are Super Data Learners

Jinjie Ni, Qian Liu, Longxu Dou, Chao Du et autres

Under strictly controlled pre-training settings, we observe a Crossover: when unique data is limited, diffusion language models (DLMs) consistently surpass autoregressive (AR) models by training for more epochs. The crossover shifts later with more or higher-quality data, earlier with larger models, and …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

Training Optimal Large Diffusion Language Models

Jinjie Ni, Qian Liu, Chao Du, Longxu Dou et autres

We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

Unnatural Languages Are Not Bugs but Features for LLMs

Keyu Duan, Yiran Zhao, Zhili Feng, Jinjie Ni et autres

Large Language Models (LLMs) have been observed to process non-human-readable text sequences, such as jailbreak prompts, often viewed as a bug for aligned LLMs. In this work, we present a systematic investigation challenging this perception, demonstrating that unnatural languages - strings that …

0 citations arXiv (Cornell University)
Accès ouvert 2025 conference-paper OpenAlex

NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation

Xiangyan Liu, Jinjie Ni, Zijian Wu, Chao Du et autres

Recent advances in reinforcement learning (RL) have strengthened the reasoning capabilities of vision-language models (VLMs). However, enhancing policy exploration to better scale test-time compute remains largely underexplored. In addition, VLMs continue to struggle with imperfect visual perception, which in turn affects the …

sg, cn (code pays fourni par la source)

0 citations
Accès ouvert 2024 preprint OpenAlex

Boosting LLM via Learning from Data Iteratively and Selectively

Qi Hui Jia, Siyu Ren, Ziheng Qin, Fuzhao Xue et autres

Datasets nowadays are generally constructed from multiple sources and using different synthetic techniques, making data de-noising and de-duplication crucial before being used for post-training. In this work, we propose to perform instruction tuning by iterative data selection (\ApproachName{}). We measure the quality …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures

Jinjie Ni, Yifan Song, Deepanway Ghosal, Bo Li et autres

Perceiving and generating diverse modalities are crucial for AI models to effectively learn from and engage with real-world signals, necessitating reliable evaluations for their development. We identify two major issues in current evaluations: (1) inconsistent standards, shaped by different communities with varying …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures

Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng et autres

Evaluating large language models (LLMs) is challenging. Traditional ground-truth-based benchmarks fail to capture the comprehensiveness and nuance of real-world queries, while LLM-as-judge benchmarks suffer from grading biases and limited query quantity. Both of them may also become contaminated over time. User-facing evaluation, …

2 citations arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.