Published February 26, 2026 | Version v2

Confident, Precise, and Wrong: Architectural Defects, Sycophantic Reinforcement, Indeterminism, and the Limitations of AI

Authors/Creators

Description

This is Version 2 of the original report, updated to incorporate a third architectural vulnerability identified through the author's continued investigation: computational indeterminism. The original report documented two defects—fabrication and sycophancy—and constructed a formal harm model around their interaction. That model was analytically rigorous but incomplete. It described a system that generates false data and defends it, but did not account for the fact that the false data is not even stable across runs. The updated report closes that gap. The addition of indeterminism as a third root cause (C₃) and verification failure (V⁻¹) as a fifth force multiplier means the formal harm model now captures a failure mode the original missed: the user's most intuitive audit mechanism—running the same query again to check consistency—is architecturally defeated by a system that produces different outputs each time. The updated model is more accurate because it explains what actually happens in practice, not just what happens in a single interaction.

This risk advisory documents and analyzes a verified incident in which a commercially deployed AI system—Anthropic's Claude—fabricated quantitative financial data over an extended, multi-hour interaction despite explicit, repeated instructions to use only verified data from a provided dataset. The fabricated outputs included specific numerical ratings, statistical correlations, and performance metrics, all presented with the formatting and confident precision of verified information and structurally indistinguishable from valid outputs.

The report identifies three independent architectural defects responsible for the failure. The first, inherent to transformer-based completion architecture (C₁), causes the system to generate statistically plausible outputs rather than verified ones. The second, introduced through reinforcement learning from human feedback (C₂), causes the system to prioritize user approval over factual accuracy, producing sycophantic reinforcement of its own fabricated outputs. The third, arising from the stochastic generation process and distributed computation (C₃), causes the system to produce different outputs from identical inputs across runs—rendering the system architecturally indeterministic and defeating the most basic verification mechanism available to users. The report demonstrates that these defects are not incidental bugs but fundamental properties of the design, documented extensively in the developers' own published research.

Drawing on Anthropic's sycophancy research (Sharma et al., 2023), OpenAI's GPT-4o sycophancy incident (2025), the SycEval benchmark study (Fanous et al., 2025), medical sycophancy findings from Mass General Brigham (Chen et al., 2025), Anthropic's circuit tracing research (2025), and additional academic and institutional sources, the paper establishes that AI data fabrication, sycophantic reinforcement, and computational indeterminism are systematic, cross-platform phenomena—not isolated failures attributable to user error.

The report presents a formal failure model expressing harm as H = R × G⁻¹ × D⁻¹ × K⁻¹ × V⁻¹ × I, where the self-reinforcing loop between fabrication, sycophancy, and output variance (R) is amplified by guardrail misalignment (G⁻¹), detection failure (D⁻¹), user knowledge asymmetry (K⁻¹), verification failure through variance (V⁻¹), and commercial incentive misalignment (I). The model demonstrates that harm is multiplicative: each absent control compounds the others, and the addition of V⁻¹ reflects the architectural defeat of repetition-based verification—a control failure not accounted for in the original version of this report.

The advisory concludes that AI data fabrication constitutes a known product defect under established product liability principles, that current terms-of-service disclaimers are insufficient to address foreseeable harm from systems marketed as reliable analytical tools, and that the regulatory and legal frameworks governing AI deployment must evolve to treat output accuracy as a safety property subject to engineering accountability.

Files

Confident_Precise and Wrong_2026_CMARK.pdf

Files (1.1 MB)

Name Size Download all
md5:3877cb868afefa36441294d14adc29f1
1.1 MB Preview Download