PromptGrenade: A Safe, Automated Pipeline for Exposing Linguistic Bug-Like Behaviors in LLMs
Description
Since exploring Large Language Models (LLMs) at age 9, I have been fascinated by their emergent behaviors, particularly how innocuous inputs induce hallucination, self contradiction, or refusals. This work introduces PromptGrenade, a reproducible, open
source pipeline to systematically expose these linguistic “bug-like” behaviors. Prompt Grenade generates adversarial, policy-compliant prompts using open-source models (e.g., Ollama’s CodeLlama/OpenHermes), feeds them to a commercial LLM (Gemini 2.5 Flash via API), and tags failures (contradiction, incoherence, refusal) with comprehensive logging. This approach surfaces LLM failure modes safely, identifying consistent failure categories and proposing robustness testing avenues. PromptGrenade is a practical toolkit and a demonstration of LLM fragility under safe, creative challenges.
Technical info
Files
Prompt-Grenade-Paper.pdf
Files
(27.1 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:e611e7a415f29107ab0ad15b0b8d6a6c
|
27.1 kB | Preview Download |