Published September 24, 2026 | Version v1.0

Fixed N & S IDs for Auditable Embedding Trace and Trackable Output N IDs - Small Experiment

Description

Problem:
Transformer is called black box because we cannot see which embedding weights were touched by input and which output IDs were top.

Idea like tokenizer:
In tokenizer, word_12 gets fixed ID 12. During training embedding values update but ID 12 stays 12. Index fixed, value updates.

Same for neurons and weights in this experiment:
- Each neuron gets fixed N ID in sequence: N0,N1,N2... In experiment 4096 neurons -> N IDs 0 to 4095 fixed.
- Each weight gets fixed S ID in sequence: S0,S1,S2... In experiment 1,179,648 weights -> S IDs 0 to 1179647 fixed.
- During training values update but N and S ID numbers stay same. S_map example: N_embedding S 0-524287, W_Q S 524288-540671, W_K S 540672-557055 etc. Counting done once, IDs fixed.

What is actually auditable and proven in fresh experiment:

1. Embedding trace has clean list — auditable:
Input ["word_12","word_45","word_78","word_200"] -> N IDs [12,45,78,200]
N 12 uses S 1536-1663 only (128 weights)
N 45 uses S 5760-5887 only
N 78 uses S 9984-10111 only
N 200 uses S 25600-25727 only
Not all weights, only these. This is logged and repeatable.

2. Output N IDs trackable:
For query, activated N IDs top10 = [2895,209,3716,1615,3982,1957,2305,478,2521,2506]
Same query again gives same list -> first == second True -> fixed and trackable. This uses threshold top-k which is standard in interpretability research.

3. Training keeps IDs fixed:
Before training N count 4096 after 4096 same, S count 1179648 after 1179647 same -> IDs fixed while values updated. Proven in fresh_experiment_proof.txt.

What is NOT claimed in this version:
Middle layers W_Q,W_K,W_V,W_O,FFN are dense — almost every weight gets small continuous number, not simple on/off. So I do not claim full middle chain in simple ordered list. Full causal trace needs more work like activation patching.

Answer to "just renaming indices" feedback:
Yes S IDs are renaming flattened indices, like tokenizer IDs are renaming. But fixed address is needed to log and audit. Without fixed ID you cannot say which S IDs were touched.

Answer to "model cannot inspect own weights" feedback:
That is true for large production models, but not for small owned model where weights are accessible. This experiment runs on owned model with full weight access.

Proof files included:
fresh_experiment_proof.txt and auditable_trace.json show fixed N and S counts, embedding S trace, top10 N repeatability.

Conclusion:
Giving fixed N ID to each neuron and fixed S ID to each weight makes embedding touch auditable and output N top10 trackable. Works for thousands or lakhs of neurons as same counting method. Full middle layer dense trace is future work.

Author: Manish Kumar Parihar

Files

Paper-Transformer-Blackbox-Final-Clear.md

Files (6.7 kB)

Name Size Download all
md5:ea894ba77a866aaf425718e23342e6e2
1.2 kB Preview Download
md5:ad5941fdf527ecbf2c2d1b37e8bb75b4
926 Bytes Preview Download
md5:29e93e4fa5c34e6f3750d9a334429426
4.6 kB Preview Download