Aller au contenu principal
Accès ouvert déclaré 2026 preprint

Adaptive Deceptive Response Framework for Prompt Injection Attacks on Large Language Models

0Citations signalées, ce qui n’est pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés

Le résumé fourni par la source

Large language models (LLMs) are increasingly deployed behind conversational and API-driven interfaces that expose them directly to untrusted natural-language input, making them susceptible to prompt injection, jailbreaking, and data-exfiltration attacks that conventional Web Application Firewalls (WAFs) are not designed to detect, since such attacks are semantically rather than syntactically anoma- lous. This paper presents a self-hosted security middleware that sits between client applications and a locally hosted LLaMA-3 model (served via Ollama) and performs multi-layer threat detection on every incoming prompt. The pipeline combines (i) leetspeak-aware keyword matching, (ii) Shannon-entropy based obfuscation detection, and (iii) LLM-based semantic classification, whose outputs are fused into a single weighted risk score that routes each request into one of three operating modes: SAFE, MONITOR, or DECEPTION. Unlike systems that merely block or warn, the DECEPTION mode actively engages suspected attackers with dynamically generated, plausible-looking fake credentials, database schemas, and system configuration data, functioning as a natural-language honeypot that wastes attacker effort while logging attacker behavior. The system further maintains a TTL-based semantic cache to reduce redundant analysis, a per-key sliding-window session store to detect multi-turn attacks, API-key authen- tication with rate limiting, an output-side data-loss-prevention (DLP) scanner that redacts personally identifiable information (PII) before responses leave the system, and a real-time monitoring dashboard. Every component is open source and runs entirely on local infrastructure, with no dependency on ex- ternal paid APIs. We describe the system’s architecture, scoring formula, and deception mechanism, and report on functional and adversarial testing performed against the implementation. We position this work as a practical, reproducible, and cost-free alternative to commercial LLM guardrail products, and discuss its current limitations and directions for quantitative evaluation.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Les sujets associés

Web Application Security VulnerabilitiesSpam and Phishing DetectionSecurity and Verification in Computing

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.