Published July 23, 2026 | Version v4

Mechanistic Validity: A Theory of Validity for Mechanistic Claims

Authors/Creators

Description

Mechanistic claims are ubiquitous in empirical science—this circuit implements this computation, this gene causes this disease, this pathway mediates this outcome. What makes such a claim valid?

We argue that a mechanistic claim is only as strong as its weakest validity dimension. Five validity types—construct, measurement, internal, external, and interpretive—form a dependency chain: construct validity is a ceiling on all others, so no amount of causal evidence rescues a claim whose target concept is incoherent. This is a theory of validity, not a taxonomy: it predicts that claims will cap at the tier corresponding to their weakest dimension, and that single high-scoring metrics cannot compensate for untested dimensions.

We operationalize this theory as Mechanistic Validity, drawing on philosophy of science, measurement theory, and validation methodology from neuroscience, pharmacology, and genetics. The framework evaluates; the methodology produces what is evaluated. For practitioners, the realistic ask is three questions: Is the description mode declared? Is the claimed tier supported by evidence? Does the evidence tier match the claim tier?

We audit fifteen mechanistic claims: sparse-autoencoder features, probing classifiers, steering vectors, and the global workspace hypothesis. The framework discriminates: the IOI circuit reaches Mechanistically Supported but caps there because cross-task specificity is untested; induction heads reach Triangulated; the gender-bias circuit is Disconfirmed. We propose evidence cards as a standard reporting unit for mechanistic claims, making gaps visible and the path to closing them explicit.

Files

mechanistic_validity_v8.pdf

Files (613.0 kB)

Name Size Download all
md5:93e2148fd71731aea2f3af31b553686b
613.0 kB Preview Download