Safeguarding Language Models via Self-Destruct Trapdoor
Rattachement africain : il. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
The potential misuse and misalignment of language models (LMs) is a central safety concern. This work presents Self-Destruct, a novel mechanism to restrict specific behaviors in LMs by leveraging overlooked properties of the underlying hardware. We observe that the LM frameworks use limited-precision formats (e.g., FP32), which are vulnerable to overflow errors during matrix multiplications. Exploiting this property, Self-Destruct replaces selected weights in pre-trained LM layers with values that act as traps, triggering a system error only when the model engages in targeted behaviors, such as harmful text generation, while leaving normal functionality unaffected. Unlike post-hoc filters, this safeguard is embedded directly within the model, introduces neither inference overhead nor auxiliary models, and requires only a set of examples for calibration. Experiments on Llama-3.2 and Qwen-2.5 demonstrate that Self-Destruct provides competitive protection against jailbreak attacks while preserving accuracy on standard benchmarks. Our results highlight the potential of hardware-aware safeguards as an efficient, low-overhead complement to existing LM defenses.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Où se fait cette recherche
-
Tel Aviv University pays non établi dans la noticeUniversité ou école supérieure
-
TAU pays non établi dans la noticeInstitution
Tel Aviv University et TAU.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.