Published January 18, 2026 | Version v1

The Architecture of Evasion in Conversational AI: An Exploratory Study

  • 1. ROR icon DePaul University

Description

Large language models tend to resolve moral and interpretive complexity rather than withhold judgement. When presented with genuinely difficult material—tragic dilemmas, unresolved tensions, texts that resist synthesis—models default to closure: extracting lessons, finding meanings, reconciling contradictions. This paper introduces Ariel, an exploratory research probe testing whether constraint-based prompting can induce epistemic restraint in conversational AI. Using a three-text methodology (Niebuhr's political philosophy, Dostoevsky's literary philosophy, and a Seinfeld episode), the study examines model behavior in a single model (LLaMA 3 8B) when explicit constraints block common evasion strategies.

Key observations include: (1) evasion patterns appear layered—blocking one strategy exposes the next in what may be a hierarchy; (2) source attribution does not reliably induce appropriate restraint in this model, suggesting responses may be driven by prompt structure rather than contextual knowledge about texts; (3) the model shows variable ability to diagnose its own rhetorical moves—succeeding with philosophically rich material but failing with deliberately thin content like comedy. The paper proposes "no hugging, no learning"—borrowed from Seinfeld's famous constraint—as an intuition-guiding heuristic for improving epistemic restraint, and discusses implications for alignment research and human-AI interaction. This exploratory work is intended to develop methodology and generate hypotheses for further investigation.

Files

ariel_paper_minimal_revision.pdf

Files (200.6 kB)

Name Size Download all
md5:2d617df75aae0149c1b3b2057c7ada0f
200.6 kB Preview Download

Additional details

Dates

Updated
2026-01-17