Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection
Description
Every open-source prompt injection detector today relies on the same two signals: surface-level pattern matching and fine-tuned ML classifiers. Both have demonstrated failure modes — a joint study by researchers from OpenAI, Anthropic, and Google DeepMind (ICLR 2025) bypassed 12 published defenses with >90% attack success rate. We propose seven novel detection techniques borrowed from forensic linguistics (stylometric discontinuity), materials science (adversarial fatigue tracking), network security (honeypot tool definitions), bioinformatics (Smith-Waterman sequence alignment), economics (prediction market ensemble), signal processing (perplexity spectral analysis), and compiler theory (taint tracking). Each technique analyzes a fundamentally different signal than existing methods. To our knowledge, none have been previously applied to prompt injection detection. All implementations are open-source (Apache 2.0) within the prompt-shield framework.
Files
cross-domain-techniques-research-paper.pdf
Files
(291.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:5812d0ccc6fc21d3f2d2acbc8ef6d9d9
|
291.4 kB | Preview Download |
Additional details
Related works
- Is supplemented by
- Software: https://github.com/mthamil107/prompt-shield (URL)