Accès ouvert
2026
preprint
OpenAlex
Bilgehan Sel, Vaishakh Keshava, Phillip Wallis, Lukas Rutishauser et autres
Addressing the critical need for robust safety in Large Language Models (LLMs), particularly against adversarial attacks and in-distribution errors, we introduce Reinforcement Learning with Backtracking Feedback (RLBF). This framework advances upon prior methods, such as BSAFE, by primarily leveraging a Reinforcement Learning …
Accès ouvert
2026
preprint
OpenAlex
Bilgehan Sel, Vaishakh Keshava, Phillip Wallis, Lukas Rutishauser et autres
Addressing the critical need for robust safety in Large Language Models (LLMs), particularly against adversarial attacks and in-distribution errors, we introduce Reinforcement Learning with Backtracking Feedback (RLBF). This framework advances upon prior methods, such as BSAFE, by primarily leveraging a Reinforcement Learning …
us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Minbeom Kim, Mihir Parmar, Phillip Wallis, Lesly Miculicich et autres
AI agents equipped with tool-calling capabilities are susceptible to Indirect Prompt Injection (IPI) attacks. In this attack scenario, malicious commands hidden within untrusted content trick the agent into performing unauthorized actions. Existing defenses can reduce attack success but often suffer from the …
Accès ouvert
2026
preprint
OpenAlex
Minbeom Kim, Mihir Parmar, Phillip Wallis, Lesly Miculicich et autres
AI agents equipped with tool-calling capabilities are susceptible to Indirect Prompt Injection (IPI) attacks. In this attack scenario, malicious commands hidden within untrusted content trick the agent into performing unauthorized actions. Existing defenses can reduce attack success but often suffer from the …
Accès ouvert
2025
preprint
OpenAlex
Zizhao Wang, Dingcheng Li, Vaishakh Keshava, Phillip Wallis et autres
Large Language Model (LLM) agents can leverage tools such as Google Search to complete complex tasks. However, this tool usage introduces the risk of indirect prompt injections, where malicious instructions hidden in tool outputs can manipulate the agent, posing security risks like …
Accès ouvert
2025
preprint
OpenAlex
Bilgehan Sel, Dingcheng Li, Phillip Wallis, Vaishakh Keshava et autres
Large language models (LLMs) have demonstrated remarkable capabilities across various tasks, but ensuring their safety and alignment with human values remains crucial. Current safety alignment methods, such as supervised fine-tuning and reinforcement learning-based approaches, can exhibit vulnerabilities to adversarial attacks and often …
2022
conference-paper
OpenAlex
Phillip Wallis, Xubo B. Song
It’s commonplace in modern deep learning to achieve SOTA performance by fine-tuning a large, pretrained base model. Recent successes in natural language processing, attributed in part to knowledge transfer from large, pretrained, transformer-based language models, have sparked a similar revolution in computer …
us
(code pays fourni par la source)
Accès ouvert
2021
preprint
OpenAlex
J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et autres
An important paradigm of natural language processing consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains. As we pre-train larger models, full fine-tuning, which retrains all model parameters, becomes less feasible. Using GPT-3 175B as an …
2020
conference-paper
OpenAlex
Phillip Wallis, Daniel B. Yaeger, Alexander B. Kain, Xubo B. Song et autres
Rapid eye movement (REM) sleep behavior disorder (RBD) is a sleep disorder that features loss of atonia, or REM sleep without atonia (RSWA). RBD and RSWA are early manifestations of degenerative neurological diseases such as Parkinson’s and Lewy Body Dementia. Accurate diagnosis …
us
(code pays fourni par la source)
Accès ouvert
2020
conference-paper
OpenAlex
MohamadAli Torkamani, Shiv Shankar Shankar, Amirmohammad Rooshenas, Phillip Wallis
Most deep neural networks use simple, fixed activation functions, such as sigmoids or rectified linear units, regardless of domain or network structure. We introduce differential equation units (DEUs), an improvement to modern neural networks, which enables each neuron to learn a particular …
de, us, gb
(code pays fourni par la source)
Accès ouvert
2020
article
OpenAlex
MohamadAli Torkamani, Amirmohammad Rooshenas, Phillip Wallis
Most deep neural networks use simple, fixed activation functions, such as sigmoids or rectified linear units, regardless of domain or network structure. We introduce differential equation units (DEUs), an improvement to modern neural networks, which enables each neuron to learn a particular …
de, us
(code pays fourni par la source)
2020
article
OpenAlex
Moragh Mackay, Ross Colliver, Phillip Wallis, Ray L. Ison et autres