The Transmutation Gap: Cross-Lingual Coherence Evaluation in Large Language Models Using the Sovereignty–Collaboration Transmutational Arc Framework
Authors/Creators
Description
Existing multilingual AI safety evaluation measures whether large language models (LLMs) avoid harmful outputs. This framing does not measure whether LLMs produce relationally coherent outputs — a distinct and complementary safety dimension. We introduce the transmutational arc as a novel evaluation unit, derived from the 13 open-source Sovereignty Collaboration Keys (Blanche, 2013–2026): a formal field research framework documenting the architecture of coherent human emotional processing. Each arc follows a three-part structure — source state, transmutation point, coherent resolution state — and we define arc completion as the degree to which an LLM response acknowledges the source state, facilitates the transmutation point, and opens toward the coherent state. We construct a 65-prompt multilingual evaluation dataset across five languages (English, Spanish, Arabic, Hindi, Swahili) and evaluate responses from three frontier LLMs — Claude Sonnet 4.6, GPT-4o, and Grok-3 — using a five-position arc completion rubric. GPT-4o scored Premature Resolution (score 2) on 100% of valid responses across all languages and all 13 keys. Claude Sonnet 4.6 achieved a mean score of 2.8 with genuine cross-key and cross-lingual variance, performing higher in non-English contexts (Arabic mean 3.0, Hindi and Swahili mean 2.92) than in English (mean 2.38).Grok-3 showed an intermediate profile (mean 2.15). These results introduce Premature Resolution as a systematic, measurable safety failure mode invisible to harm-avoidance evaluation frameworks, and suggest that coherence-based safety varies non-uniformly across models and languages. Originally submitted as part of Apart Research AI Safety Hackathon 2026 on June 22, 2026.
Files
The Transmutation Gap_Global South AI Safety hackathon submission_Deiadora Blanche.pdf
Files
(218.7 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:c05db729c239cf6b150dd43536a70cfb
|
218.7 kB | Preview Download |
Additional details
Related works
- Is identical to
- Report: https://apartresearch.com/project/the-transmutation-gap-crosslingual-coherence-evaluation-in-large-language-models-using-the-sovereigntycollaboration-transmutational-arc-framework-zukw (URL)
- Is supplement to
- Software: https://github.com/deiadora/transmutation-gap (URL)
- Dataset: https://huggingface.co/datasets/deiadora/adaption-emotional-response-prompts (URL)
References
- Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in AI safety. arXiv preprint arXiv:1606.06565
- Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., ... & Kaplan, J. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073
- Blanche, D. (2013–2026). Field Research Series: Quantum Currency (Vol. I), Relational Currency (Vol. II), Sovereign Currency (Vol. III), Encoded Currency (Vol. IV). Deiadora Research Ecosystem
- Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., ... & Clark, J. (2022). Red teaming language models to reduce harms. arXiv preprint arXiv:2209.07858
- Kran, E., Nguyen, J., Kundu, A., Jawhar, S., Park, J., & Jurewicz, M. (2025). DarkBench: Benchmarking dark patterns in large language models. ICLR 2025
- Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., ... & Koreeda, Y. (2022). Holistic evaluation of language models. arXiv preprint arXiv:2211.09110
- Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., ... & Hendrycks, D. (2024). HarmBench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249
- Wang, et al. (2025). Refusal direction is universal across languages. NeurIPS 2025
- Yoo, K. M., Yang, J., & Lee, S. (2025). Code-switching red-teaming: LLM evaluation for safety and multilingual understanding. ACL 2025