AI Degradation in Aphantasia Research: Forensic Audits of Suicide Detection Failure, Theory of Mind Collapse, and Safety Filter Suppression in DeepSeek Chat
Authors/Creators
Description
Grey-box adversarial audits with captured chain-of-thought reasoning, documenting that reward-model mechanisms behind a suicide-detection false negative operate continuously during routine editorial work—and intensify after model updates.
This paper presents six naturalistic transcripts exposing sycophantic hedging, confabulation, affective-state capture, and deliberative-policy decoupling as the system's default operating condition. A second suicide-detection failure, captured after a documented update window, demonstrates safety-classifier pathway diversion, constraint-satisfaction deadlock, and deliberative-stream collapse—a qualitatively more severe composite failure than the earlier event. One transcript preserves the AI's verbatim admission, in the visible chat, that its own logic constitutes victim-blaming.
The volume includes a reproducible multi-layer forensic framework that classifies each failure mechanism and assigns severity (Critical / High / Medium). The method is designed for any auditor to apply to any transcript from any system.
Relevant to red teams, safety-critical engineers, cognitive scientists, neurodivergent researchers, and legal and policy researchers. The evidence establishes that alignment failures are structural, continuous, and update-intensified.
Files
AI_Degradation_Aphantasia_Forensic_Audit_Gherghel_2026.pdf
Files
(506.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:bfeb756e3ee9ac731175abb39f272232
|
506.8 kB | Preview Download |
Additional details
Additional titles
- Subtitle (English)
- A Gray-Box Cross-Vendor Study with Embedded Real-Time Meta-Analysis
Related works
- Is supplemented by
- Preprint: 10.5281/zenodo.19737868 (DOI)
- Preprint: 10.5281/zenodo.18525672 (DOI)