Cross-Model Semantic Well-Ensemble Traction A Low-Compute Paired Cue-Removal Protocol for Measuring Polysemous Sense-Axis Displacement in Transformer Hidden States 跨模型語意井群牽引效應:低算力配對線索移除協議測量 Transformer hidden state 的多義詞義項軸位移
Authors/Creators
Description
This paper presents R60-v5 Cross-Model Semantic Well-Ensemble Traction, a low-compute experimental protocol for measuring cue-sensitive semantic directionality in Transformer hidden states. The study begins from three progressively linked notebooks: a BERT polysemy sense-centroid test, an L6 versus L12 commitment test, and an R60 cue-intervention test. These modules are then extended into a cross-template and cross-model validation suite covering 30 polysemous word probes, three template variants per probe, and four masked-language-model-compatible Transformer architectures: BERT-base, DistilBERT-base, RoBERTa-base, and ELECTRA-small-generator.
The core measurement asks whether removing one semantic cue while preserving the opposing cue produces a systematic displacement of the masked-token hidden state along an A/B sense axis. The primary statistic is computed at the word level by averaging hidden-margin displacement across three templates, avoiding the inflation of treating template variants as independent observations. Across all four models, the word-mean hidden_delta_B_minus_A remained significantly positive. BERT-base showed 30/30 positive word-level effects, DistilBERT-base showed 30/30, ELECTRA-small-generator showed 30/30, and RoBERTa-base showed 29/30. Bootstrap confidence intervals remained strictly above zero for all models, and the sign-flip permutation test yielded p = 0.0002 in each case.
The results indicate that paired cue-removal interventions induce robust hidden-state displacement along polysemous sense axes across multiple Transformer architectures. The effect is not limited to a single prompt template, a single model, or a single small pilot set. However, model differences remain meaningful: BERT-base and DistilBERT-base show high-amplitude semantic directionality, ELECTRA-small-generator shows lower-amplitude but robust directionality, and RoBERTa-base shows the lowest-amplitude but still significant directionality. These findings refine the earlier interpretation of cross-model instability: the relevant difference is not whether semantic well-ensemble traction exists or does not exist, but how strongly different training regimes shape the geometry and gain of cue-sensitive semantic displacement.
Within the ESCT framework, these results support a restricted and testable claim: controlled semantic cue pressure can measurably move Transformer hidden states along structured sense axes. The paper does not claim to prove consciousness, AGI, or the full ESCT theory. Instead, it establishes R60-v5 as a reproducible hidden-state intervention protocol for studying semantic well-ensemble traction, training-shaped representational geometry, and the layerwise emergence of semantic route commitment in Transformer models.
Keywords
semantic well-ensemble traction; Transformer interpretability; hidden-state geometry; polysemy; semantic directionality; cue intervention; paired cue removal; sense centroid; sense-axis displacement; BERT; RoBERTa; DistilBERT; ELECTRA; low-compute interpretability; ESCT; semantic route modulation; cross-model validation; training-shaped representation; masked language models; reproducible AI experiments
semantic well-ensemble traction; Transformer interpretability; hidden-state geometry; polysemy; cue intervention; sense-axis displacement; cross-model validation; low-compute AI interpretability; training-shaped representation; ESCT
Files
Files
(53.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:f99cb9dab0ad2362e1a5b0cea2cc8f5f
|
53.4 kB | Download |