Accès ouvert
2025
preprint
OpenAlex
Lama Alssum, Hani Itani, Hasan Abed Al Kader Hammoud, Philip H. S. Torr et autres
The safety alignment of large language models (LLMs) is becoming increasingly important with their democratization. In this paper, we study the safety degradation that comes with adapting LLMs to new tasks. We attribute this safety compromise to catastrophic forgetting and frame the …
Accès ouvert
2025
preprint
OpenAlex
Boyuan Chen, S. S. Fang, Jiaming Ji, Yanxu Zhu et autres
As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an empirically demonstrated risk across language models, AI agents, and emerging frontier systems. This project provides …
us
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Alex Dobra, Jakiw Pidstrigach, Tim Reichelt, Paolo Fraccaro et autres
Sensitivity analysis is a cornerstone of climate science, essential for understanding phenomena ranging from storm intensity to long-term climate feedbacks. However, computing these sensitivities using traditional physical models is often prohibitively expensive in terms of both computation and development time. While modern …
Accès ouvert
2025
preprint
OpenAlex
Tala Aljaafari, Varun Kanade, Philip H. S. Torr, Christian Schroeder de Witt
Deploying reinforcement learning (RL) in safety-critical settings is constrained by brittleness under distribution shift. We study out-of-distribution (OOD) detection for RL time series and introduce DEEDEE, a two-statistic detector that revisits representation-heavy pipelines with a minimal alternative. DEEDEE uses only an episodewise …
Accès ouvert
2025
conference-paper
OpenAlex
Alejandro Ciocci Pardo, Fabio Pizzati, Tong Zhang, Alexander Pondaven et autres
Match-cuts are powerful cinematic tools that create seamless transitions between scenes, delivering strong visual and metaphorical connections. However, crafting match-cuts is a challenging, resource-intensive process requiring deliberate artistic planning. In MatchDiffusion, we present the first training-free method for match-cut generation using text-to-video …
ca, sa, ae, gb
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Tajamul Ashraf, Umair Nawaz, Abdelrahman Shaker, Rao Muhammad Anwer et autres
Vision language models (VLMs) are increasingly deployed as controllers with access to external tools for complex reasoning and decision-making, yet their effectiveness remains limited by the scarcity of high-quality multimodal trajectories and the cost of manual annotation. We address this challenge with …
Accès ouvert
2025
preprint
OpenAlex
X.D. Xue, Yifan Zhou, Guibin Zhang, Yijiang Li et autres
Self-evolution is a central research topic in enabling large language model (LLM)-based agents to continually improve their capabilities after pretraining. Recent research has witnessed a transition from reinforcement learning (RL)-free to RL-based methods. Current RL-based methods either rely on dense external reward …
Accès ouvert
2025
preprint
OpenAlex
S. Chang, Junchi Yu, Weixing Wang, Yongqiang Chen et autres
Diffusion large language models (D-LLMs) have recently emerged as a promising alternative to auto-regressive LLMs (AR-LLMs). However, the hallucination problem in D-LLMs remains underexplored, limiting their reliability in real-world applications. Existing hallucination detection methods are designed for AR-LLMs and rely on signals …
Accès ouvert
2025
preprint
OpenAlex
Fengyuan Liu, Rui Zhao, Shuo Chen, Guohao Li et autres
Individual Large Language Models (LLMs) have demonstrated significant capabilities across various domains, such as healthcare and law. Recent studies also show that coordinated multi-agent systems exhibit enhanced decision-making and reasoning abilities through collaboration. However, due to the vulnerabilities of individual LLMs and …
Accès ouvert
2025
preprint
OpenAlex
Josefa Lia Stoisser, Marc Boubnovski Martell, Lawrence Phillips, Gianluca Mazzoni et autres
Large language model (LLM) agents are increasingly deployed in structured biomedical data environments, yet they often produce fluent but overconfident outputs when reasoning over complex multi-table data. We introduce an uncertainty-aware agent for query-conditioned multi-table summarization that leverages two complementary signals: (i) …
Accès ouvert
2025
article
OpenAlex
Yangchen Pan, Junfeng Wen, Chenjun Xiao, Philip H. S. Torr
Background: Traditional supervised learning (SL) assumes data points are independently and identically distributed (i.i.d.), which overlooks dependencies in real-world data. Reinforcement learning (RL), in contrast, models dependencies through state transitions. Objectives: This study aims to bridge SL and RL by reformulating SL …
gb, ca
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Jin Myung Kwak, Lama Alssum, Bernard Ghanem, Philip H. S. Torr et autres
Fine-tuning language models is commonly believed to inevitably harm their safety, i.e., refusing to respond to harmful user requests, even when using harmless datasets, thus requiring additional safety measures. We challenge this belief through systematic testing, showing that poor optimization choices, rather …