AEGIS: Real-Time Latent-Space Backdoor Detection for Dependable and Secure Small Language Model Inference
Résumé fourni par la source
Backdoor attacks pose a serious threat to small language models (SLMs) because compromised models can behave normally on benign inputs while producing attacker-specified outputs when a hidden trigger is activated. Existing defenses commonly require model retraining, operate only before deployment, rely on input-level signals, or incur excessive inference overhead. This paper presents AEGIS (Activation Evaluation and Guardrail for Inference Security), a non-invasive runtime framework for detecting backdoor-induced anomalies in transformer representations. AEGIS dynamically monitors equidistant internal layers, mean-pools and concatenates their hidden states, and compresses the resulting high-dimensional representation into a 64-dimensional latent space. A hybrid compression mechanism uses zero-shot principal component analysis for predominantly linear text representations and a lightweight autoencoder for nonlinear or multimodal representations. Detection is then performed entirely on the GPU using cosine distance from a centroid calibrated on clean data, without modifying or retraining the protected model. We evaluate AEGIS across Mistral-7B, Qwen2.5-7B, a 4-bit QLoRA-backdoored Qwen2.5-1.5B model, and a BadNets-backdoored Vision Transformer, covering simulated latent and PEFT-style attacks, genuine fine-tuned backdoors, FP16 inference, and 4-bit NF4 quantization. Across seven experimental scenarios, AEGIS achieves AUROC values between 0.95 and 1.00, true-positive rates between 86% and 100%, and detection latency between 0.47 and 2.40 ms. It obtains AUROC = 1.00 and 100% true-positive rate against the genuinely fine-tuned QLoRA backdoor, while achieving AUROC=0.997 against the fine-tuned visual BadNets attack. Under 4-bit quantization, it retains AUROC=0.95 with 1.50 ms latency and a 4.52 GB memory footprint. These results demonstrate that GPU-native latent-space monitoring can provide effective, retraining-free backdoor detection for real-time transformer inference.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- AEGIS: Real-Time Latent-Space Backdoor Detection for Dependable and Secure Small Language Model Inference
- Date Crossref
- 31/08/2026
- Éditeur
- MDPI AG
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.