Published November 21, 2025 | Version v1

BlockCert: Certified Blockwise Extraction of Transformer Mechanisms

Authors/Creators

Description

Mechanistic interpretability aspires to reverse-engineer neural networks into explicit algo-
rithms, while model editing seeks to modify specific behaviours without retraining. Both areas
are typically evaluated with informal evidence and ad-hoc experiments, with few explicit guar-
antees about how far an extracted or edited model can drift from the original on relevant
inputs. We introduce BlockCert, a framework for certified blockwise extraction of transformer
mechanisms, and outline how a lightweight extension can support certified local edits. Given a
pre-trained transformer and a prompt distribution, BlockCert extracts structured surrogate
implementations for residual blocks together with machine-checkable certificates that bound
approximation error, record coverage metrics, and hash the underlying artifacts. We formalize a
simple Lipschitz-based composition theorem in Lean 4 that lifts these local guarantees to a global
deviation bound. Empirically, we apply the framework to GPT-2 small, TinyLlama-1.1B-Chat,
and Llama-3.2-3B. Across these models we obtain high per-block coverage and small residual
errors on the evaluated prompts, and in the TinyLlama setting we show that a fully stitched
model matches the baseline perplexity within ≈6 ×10−5on stress prompts. Our results suggest
that blockwise extraction with explicit certificates is feasible for real transformer language models
and offers a practical bridge between mechanistic interpretability and formal reasoning about
model behaviour

Files

blockcert_arxiv.pdf

Files (444.2 kB)

Name Size
md5:67061951d853884a8c56874bc1f9d737
444.2 kB Preview Download