Accès ouvert
2026
preprint
OpenAlex
Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri et autres
Prompt injection is widely recognized as a major security threat to AI agents that interact with untrusted external data, such as websites, documents, and emails. Prior work has shown that, in the text domain, black-box prompt injection can achieve near-perfect attack success …
Accès ouvert
2026
preprint
OpenAlex
Jaiden Fairoze, Neal Mangaokar, Kamalika Chaudhuri, Sanjam Garg et autres
For AI agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into …
Accès ouvert
2026
preprint
OpenAlex
Pengrun Huang, Chhavi Yadav, Ruihan Wu, Kamalika Chaudhuri
Large language models (LLMs) are increasingly fine-tuned on domain-specific datasets that may contain sensitive, dataset-level properties. Recent work has shown that such dataset-level information can be effectively extracted through property inference attacks, posing a confidentiality risk. Existing defenses against these attacks primarily …
Accès ouvert
2026
preprint
OpenAlex
Pengrun Huang, Chhavi Yadav, Ruihan Wu, Kamalika Chaudhuri
Large language models (LLMs) are increasingly fine-tuned on domain-specific datasets that may contain sensitive, dataset-level properties. Recent work has shown that such dataset-level information can be effectively extracted through property inference attacks, posing a confidentiality risk. Existing defenses against these attacks primarily …
us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Mihai Christodorescu, Earlence Fernandes, Ashish Hooda, Somesh Jha et autres
We take the position that agent security must be approached as a systems problem: the AI model powering the agent must be treated as an untrusted component, and security invariants must be enforced at the system level. Through this lens, efforts to …
Accès ouvert
2026
preprint
OpenAlex
Mihai Christodorescu, Earlence Fernandes, Ashish Hooda, Somesh Jha et autres
We take the position that agent security must be approached as a systems problem: the AI model powering the agent must be treated as an untrusted component, and security invariants must be enforced at the system level. Through this lens, efforts to …
us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Cristina Menghini, Peter Ney, Hamza Kwisaba, Zifan et autres
Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framework, along with the evidence that informed our launch decision. We then discuss additional considerations, …
us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Cristina Menghini, Peter Ney, Hamza Kwisaba, Zifan et autres
Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framework, along with the evidence that informed our launch decision. We then discuss additional considerations, …
Accès ouvert
2026
preprint
OpenAlex
Pengrun Huang, Kamalika Chaudhuri, Yu-Xiang Wang
Large language models (LLMs) are pre-trained and post-trained on vast amounts of loosely curated data, raising the possibility that these models may have been trained on proprietary datasets or the same benchmarks used for evaluation. This motivates the need for dataset watermarking: …
Accès ouvert
2026
preprint
OpenAlex
Pengrun Huang, Kamalika Chaudhuri, Yu-Xiang Wang
Large language models (LLMs) are pre-trained and post-trained on vast amounts of loosely curated data, raising the possibility that these models may have been trained on proprietary datasets or the same benchmarks used for evaluation. This motivates the need for dataset watermarking: …
Accès ouvert
2026
preprint
OpenAlex
Erchi Wang, Pengrun Huang, Eli Chien, Om Thakkar et autres
Differential privacy (DP) has a wide range of applications for protecting data privacy, but designing and verifying DP algorithms requires expert-level reasoning, creating a high barrier for non-expert practitioners. Prior works either rely on specialized verification languages that demand substantial domain expertise …
Accès ouvert
2026
preprint
OpenAlex
Erchi Wang, Pengrun Huang, Eli Chien, Om Thakkar et autres
Differential privacy (DP) has a wide range of applications for protecting data privacy, but designing and verifying DP algorithms requires expert-level reasoning, creating a high barrier for non-expert practitioners. Prior works either rely on specialized verification languages that demand substantial domain expertise …
us, tw
(code pays fourni par la source)