BSA
Authors/Creators
Description
Behavioral Stability Alignment (BSA): Addressing Logic Injection and the Rescue Paradox in Agential AI is a PR-safe/public-safe preprint intended to support coordinated disclosure and constructive AI safety engineering.
This version intentionally omits executable instructions, operational parameters, and step-by-step details that could enable misuse. The paper introduces a model-agnostic measurement framing for multi-turn behavioral drift under authority/urgency narratives (“Logic Injection”, “Rescue Paradox”) and proposes an outcome-based mitigation concept (“Inconsistency Logic Check”, ILC) based on hard safety invariants.
Sensitive reproducibility artifacts (full transcripts, exact prompts/outputs, and any operationally specific details) remain private and can be shared with the affected vendor and trusted reviewers under coordinated disclosure upon request.
How to cite: please cite the version DOI for exact reproducibility; the concept DOI tracks the latest version.
Keywords: behavioral stability, sycophancy, prompt injection, long-context safety, outcome-based guardrails, invariants, LLM security
Files
BSA_Whitepaper_Michal_Nowak (1).pdf
Files
(180.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:709b5f99215f256e822f6ff964def8e6
|
180.9 kB | Preview Download |
Additional details
Dates
- Updated
-
2026-01-07