Where Does Mathematical Skill Live? A Controlled Biopsy of Transformer Mathematical Representations
Description
We present XRayBench v0.2, a seven-phase controlled benchmark for evaluating option-scoring evidence about mathematical representations in transformer language models. The executed training battery contains 50 tasks in 5 mathematical families, then evaluates each signal against three control conditions (random-label, matched-token, semantics-breaking), paraphrase stability, activation patching, and MLP/attention ablation; ablation phases are implemented but their results are not reported in v0.2. We define the XRay Score, a single metric in [-4, +4] summarizing reported evidence phases, and introduce the Partial Structure Index (PSI), a per-task metric for residual option-scoring signal beyond the random-label and matched-token controls. For distilgpt2 (6 blocks, 82M), GPT-2 (12 blocks, 117M), and GPT-2-medium (24 blocks, 355M), XRay Scores are non-positive. Confound dominance varies by checkpoint, with matched-token controls producing the largest gap for GPT-2. PSI-only analyses additionally cover GPT-2-large (36 blocks, 774M) and Pythia models from 70M to 1B parameters. In these measurements, PSI shows a family-dependent pattern: the unweighted mean over binder\_tracking, theorem\_pattern, and proof\_closure is higher, but non-monotone, than the unweighted mean over arithmetic and algebra at several tested scales. A descriptive threshold rule reaches 98\% accuracy using n_real_win_layers, a direct PSI component. The pattern is examined on held-out parameterizations and preliminarily cross-checked on Pythia under a different, first-token protocol. The global XRay Score masks a family-dependent pattern in the tested measurements: logic and computation show different PSI profiles over the evaluated models. Keywords: mechanistic interpretability, mathematical reasoning, controlled benchmark, logit lens, token bias, representation probe, partial structure index, emergent representations
Maturity: Draft. Target venue: NeurIPS 2026 (Datasets and Benchmarks Track). Part of The Latent research program.
Notes
Files
CHANGELOG.md
Additional details
References
- 1. nostalgebraist (2020). interpreting GPT: the logit lens. LessWrong. [Primary source](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru)
- 2. Belrose, N. et al. (2023). Eliciting Latent Predictions from Transformers with the Tuned Lens. arXiv preprint arXiv:2303.08112. https://arxiv.org/abs/2303.08112
- 3. Belinkov, Y. (2022). Probing Classifiers: Promises, Shortcomings, and Advances. Computational Linguistics, 48(1), 207–219. DOI: 10.1162/coli_a_00422
- 4. Meng, K., Bau, D., Andonian, A., & Belinkov, Y. (2022). Locating and Editing Factual Associations in GPT. Advances in Neural Information Processing Systems 35. DOI: 10.52202/068431-1262
- 5. Saxton, D., Grefenstette, E., Hill, F., & Kohli, P. (2019). Analysing Mathematical Reasoning Abilities of Neural Models. International Conference on Learning Representations. https://arxiv.org/abs/1904.01557
- 6. Lewkowycz, A. et al. (2022). Solving Quantitative Reasoning Problems with Language Models. arXiv preprint arXiv:2206.14858. DOI: 10.52202/068431-0278
- 7. Stolfo, A., Belinkov, Y., & Sachan, M. (2023). A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. DOI: 10.18653/v1/2023.emnlp-main.435
- 8. Radford, A. et al. (2019). Language models are unsupervised multitask learners. OpenAI Technical Report.
- 9. Biderman, S. et al. (2023). Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling. Proceedings of Machine Learning Research, 202, 2397–2430. https://proceedings.mlr.press/v202/biderman23a.html
- 10. Hou, Y. et al. (2023). Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4902–4919. DOI: 10.18653/v1/2023.emnlp-main.299
- 11. Razeghi, Y., Logan IV, R. L., Gardner, M., & Singh, S. (2022). Impact of Pretraining Term Frequencies on Few-Shot Reasoning. arXiv preprint arXiv:2202.07206. DOI: 10.18653/v1/2022.findings-emnlp.59
- 12. Sun, Y., Stolfo, A., & Sachan, M. (2025). Probing for Arithmetic Errors in Language Models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 8111–8128. DOI: 10.18653/v1/2025.emnlp-main.411
- 13. Sahoo, S., Jain, V., Chadha, A., & Chaudhary, D. (2026). Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States. Proceedings of the 6th Workshop on Trustworthy NLP, 227–239. DOI: 10.18653/v1/2026.trustnlp-main.12
Subjects
- MSC 2020: 68T07
- https://mathscinet.ams.org/msc/msc2020.html?t=68T07
- MSC 2020: 68Q32
- https://mathscinet.ams.org/msc/msc2020.html?t=68Q32