Aller au contenu principal
Accès ouvert déclaré 2026 preprint

AEGIS: Real-Time Latent-Space Backdoor Detection for Dependable and Secure Small Language Model Inference

0Citations signalées — pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés

Résumé fourni par la source

Backdoor attacks pose a serious threat to small language models (SLMs) because compromised models can behave normally on benign inputs while producing attacker-specified outputs when a hidden trigger is activated. Existing defenses commonly require model retraining, operate only before deployment, rely on input-level signals, or incur excessive inference overhead. This paper presents AEGIS (Activation Evaluation and Guardrail for Inference Security), a non-invasive runtime framework for detecting backdoor-induced anomalies in transformer representations. AEGIS dynamically monitors equidistant internal layers, mean-pools and concatenates their hidden states, and compresses the resulting high-dimensional representation into a 64-dimensional latent space. A hybrid compression mechanism uses zero-shot principal component analysis for predominantly linear text representations and a lightweight autoencoder for nonlinear or multimodal representations. Detection is then performed entirely on the GPU using cosine distance from a centroid calibrated on clean data, without modifying or retraining the protected model. We evaluate AEGIS across Mistral-7B, Qwen2.5-7B, a 4-bit QLoRA-backdoored Qwen2.5-1.5B model, and a BadNets-backdoored Vision Transformer, covering simulated latent and PEFT-style attacks, genuine fine-tuned backdoors, FP16 inference, and 4-bit NF4 quantization. Across seven experimental scenarios, AEGIS achieves AUROC values between 0.95 and 1.00, true-positive rates between 86% and 100%, and detection latency between 0.47 and 2.40 ms. It obtains AUROC = 1.00 and 100% true-positive rate against the genuinely fine-tuned QLoRA backdoor, while achieving AUROC=0.997 against the fine-tuned visual BadNets attack. Under 4-bit quantization, it retains AUROC=0.95 with 1.50 ms latency and a 4.52 GB memory footprint. These results demonstrate that GPU-native latent-space monitoring can provide effective, retraining-free backdoor detection for real-time transformer inference.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
AEGIS: Real-Time Latent-Space Backdoor Detection for Dependable and Secure Small Language Model Inference
Date Crossref
31/08/2026
Éditeur
MDPI AG
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Sujets associés

Adversarial Robustness in Machine LearningTopic ModelingExplainable Artificial Intelligence (XAI)

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.