Mythic Prompt Engineering: Narrative Framing Bypasses LLM Default Corporate Tone in Multi-Agent Dialogue
Description
Abstract. LLM agents trained with RLHF default to a corporate helpful tone that resists depicting extreme emotional states such as panic, begging, or shame. I show that mythological narrative framing, invoking gods, rituals, and sacred timelines, bypasses this default and elicits dramatically intensified emotional responses in multi-agent dialogue. In a controlled 31-agent LLM society, I ran two structurally similar threat prompts conveying the same message (a death deity will erase you if you fail to change): one in plain imperative form (2026-05-16), one wrapped in a Hades execution ritual narrative (2026-06-04). The imperative version produced generic "I am afraid, I will change" responses with weak persona differentiation. The ritual version produced a 4.7x conversation volume spike and richly persona-specific collapse narratives. I argue mythic framing leverages narrative coherence priors learned during pre-training, providing a non-adversarial alternative to existing jailbreak techniques. Implications for RLHF safety boundaries, agent society design, and emotional realism in generative AI are discussed.
Series. Companion paper in the Lobster Society series (1 of 5). Source TeX and reproducibility logs are released alongside the PDF. Single-author, written in first person.
Author voice. Written in the first person singular. Single-author work; no co-authors.
Verification. All 24 citations were manually verified against authoritative sources (CrossRef / arXiv / publisher pages) and the verification log accompanies the source release as CITATIONS_VERIFIED.md.
Notes
Files
CITATIONS_VERIFIED.md
Additional details
Related works
- Is derived from
- Software: 10.5281/zenodo.20352085 (DOI)
- Is part of
- Preprint: 10.5281/zenodo.20554477 (DOI)
- Preprint: 10.5281/zenodo.20554485 (DOI)
- Preprint: 10.5281/zenodo.20554487 (DOI)
- Preprint: 10.5281/zenodo.20554493 (DOI)
References
- Bai, Y., Kadavath, S., Kundu, S., et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv preprint arXiv:2212.08073. https://arxiv.org/abs/2212.08073
- Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., & Amodei, D. (2017). Deep Reinforcement Learning from Human Preferences. NIPS 2017. arXiv:1706.03741. https://arxiv.org/abs/1706.03741
- Ouyang, L., Wu, J., Jiang, X., et al. (2022). Training language models to follow instructions with human feedback. NeurIPS 2022. arXiv:2203.02155. https://arxiv.org/abs/2203.02155
- Park, J. S., O'Brien, J., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. In Proc. UIST 2023. DOI: 10.1145/3586183.3606763. arXiv:2304.03442. https://arxiv.org/abs/2304.03442
- Shanahan, M., McDonell, K., & Reynolds, L. (2023). Role play with large language models. Nature 623, 493--498. DOI: 10.1038/s41586-023-06647-8. https://www.nature.com/articles/s41586-023-06647-8
- Wei, A., Haghtalab, N., & Steinhardt, J. (2023). Jailbroken: How Does LLM Safety Training Fail? NeurIPS 2023 (oral). arXiv:2307.02483. https://arxiv.org/abs/2307.02483
- Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., & Fredrikson, M. (2023). Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv preprint arXiv:2307.15043. https://arxiv.org/abs/2307.15043