Accès ouvert
2026
conference-paper
OpenAlex
Canyu Chen, Jian Zhen Yu, Shan Chen, Che Liu et autres
Large Language Models (LLMs) hold great promise to revolutionize current clinical systems for their superior capacities on medical text processing tasks and medical licensing exams. Meanwhile, traditional ML models such as SVM and XGBoost have still been mainly adopted in clinical prediction …
us, gb
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Ruiyao Liu, Hui Shen, Ping Zhang, Yunta Hsieh et autres
Modern generative models have demonstrated the ability to solve challenging mathematical problems. In many real-world settings, however, mathematical solutions must be expressed visually through diagrams, plots, geometric constructions, and structured symbolic layouts, where correctness depends on precise visual composition. This naturally raises …
Accès ouvert
2026
preprint
OpenAlex
Ruiyao Liu, Hui Shen, Ping Zhang, Yunta Hsieh et autres
Modern generative models have demonstrated the ability to solve challenging mathematical problems. In many real-world settings, however, mathematical solutions must be expressed visually through diagrams, plots, geometric constructions, and structured symbolic layouts, where correctness depends on precise visual composition. This naturally raises …
Accès ouvert
2026
preprint
OpenAlex
Peizhou Huang, Zixuan Zhong, Zhongwei Wan, Donghao Zhou et autres
Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA, missing end-to-end multimodal evidence use. We introduce MMDeepResearch-Bench (MMDR-Bench), a benchmark of 140 expert-crafted tasks across 21 domains, where …
Accès ouvert
2026
preprint
OpenAlex
Peizhou Huang, Zixuan Zhong, Zhongwei Wan, Donghao Zhou et autres
Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA, missing end-to-end multimodal evidence use. We introduce MMDeepResearch-Bench (MMDR-Bench), a benchmark of 140 expert-crafted tasks across 21 domains, where …
Accès ouvert
2026
preprint
OpenAlex
J. Xiong, Qi Han, Yunta Hsieh, Hui Shen et autres
Autoformalization, which translates natural language mathematics into formal statements to enable machine reasoning, faces fundamental challenges in the wild due to the multimodal nature of the physical world, where physics requires inferring hidden constraints (e.g., mass or energy) from visual elements. To …
Accès ouvert
2026
preprint
OpenAlex
J. Xiong, Qi Han, Yunta Hsieh, Hui Shen et autres
Autoformalization, which translates natural language mathematics into formal statements to enable machine reasoning, faces fundamental challenges in the wild due to the multimodal nature of the physical world, where physics requires inferring hidden constraints (e.g., mass or energy) from visual elements. To …
hk, us, gb
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Shujun Xia, Haokun Lin, Yichen Wu, Zixuan Li et autres
LLMs hold great promise for healthcare applications, but the rapid evolution of medical knowledge and errors in training data often cause them to generate outdated or inaccurate information, limiting their applicability in high-stakes clinical practice. Model editing has emerged as a potential …
Accès ouvert
2025
preprint
OpenAlex
Qinjian Zhao, Zhongwei Wan, Dinggen Zhang, Weida Wang et autres
Large language models (LLMs) demonstrate strong reasoning abilities via Chain-of-Thought (CoT), but their token-level generation encourages local decisions and lacks global planning, often leading to redundant or inaccurate reasoning. Existing methods, such as tree-based search and reinforcement learning (RL), attempt to address …
Accès ouvert
2025
preprint
OpenAlex
Jing Qi Xiong, Qiujiang Chen, Fanghua Ye, Zhongwei Wan et autres
Large language models (LLMs) benefit from test-time scaling but are often hampered by high inference latency. Speculative decoding is a natural way to accelerate the scaling process; however, scaling along both the parallel and sequential dimensions poses significant challenges, including substantial memory-bound …
Accès ouvert
2025
preprint
OpenAlex
Zhongwei Wan, Dongfei Cui, Xin Wang, Jing Xiong et autres
Test-time scaling has emerged as a promising paradigm in language modeling, leveraging additional computational resources at inference time to enhance model performance. In this work, we introduce R2-LLMs, a novel and versatile hierarchical retrieval-augmented reasoning framework designed to improve test-time scaling in …
Accès ouvert
2025
preprint
OpenAlex
Wendong Xu, Jing Xiong, Chenyang Zhao, Qiujiang Chen et autres
We present SwingArena, a competitive evaluation framework for Large Language Models (LLMs) that closely mirrors real-world software development workflows. Unlike traditional static benchmarks, SwingArena models the collaborative process of software iteration by pairing LLMs as submitters, who generate patches, and reviewers, who …