A Lightweight Transformer Encoder for Prompt Injection Detection
Authors/Creators
Description
Prompt injection attacks pose a significant and growing threat to large language model (LLM)-based systems, enabling adversaries to override the model instructions and manipulate the output in ways that undermine security and reliability and put the whole ecosystem at stake without proper solutions. Although existing detection approaches have largely relied on LLM-based fine-tuned solutions, these methods can suffer from high computational requirements to be deployed locally in
resource-constrained environments. In response to the need to build computeefficient, accurate, and latency-driven defence mechanisms, we designed, trained, and evaluated a 2-head, 2-layer self-attention transformer encoder using prompt injection data. While our model relies on 995,970 parameters, we compared it with transformer-based models ranging from 500 million to 14 billion parameters tested in the literature on identical datasets and it demonstrated superior performance. Furthermore, the trained model achieved an average inference time less than 0.5 ms when executed on the CPU of a virtual machine equipped with 120 GB of
system memory and a single NVIDIA A100 GPU (40 GB VRAM). This demonstrates that the transformer architecture’s capabilities can be effectively utilized to create lightweight solutions specifically designed for prompt-injection detection without the need to depend on large models.
Files
MANUSCRIPT_.pdf
Files
(1.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:c4c86761571c4ac9f290e4df9f9f5ba0
|
1.2 MB | Preview Download |