Semantic Grounding and the Preservation of Information in Recursive Systems
Authors/Creators
Description
This article presents a revised information-theoretic framework for semantic grounding and recursive training. The central claim is now a rate condition rather than a fixed collapse sequence: recursive or under-corrected training introduces diagnostic degradation pressure, \(E(t)\), while independent data, verification, or other grounding channels supply effective corrective capacity, \(C_{\mathrm{eff}}(t)\). Long-term semantic viability requires that corrective capacity remain at least comparable to degradation pressure over the relevant horizon and diagnostic target.
This version formally retracts the earlier prediction that out-of-distribution or grounded accuracy should generally degrade before validation perplexity rises. Instead, the temporal lag between grounded-task degradation and perplexity degradation is treated as an empirical discriminator whose sign may vary by correction regime, model scale, benchmark sensitivity, floor effects, and grading method.
The paper publishes a preregistered evaluation design for testing this revised claim across five recursive-training regimes, three seeds, and two Qwen2.5 model scales. The expected signal is comparative: low-correction regimes should show stronger diagnostic separation or instability than regimes with fresher or higher-quality correction. The preregistered empirical evaluation itself is reserved for a later version.
Files
Semantic Grounding and the Preservation of Information in Recursive Systems (10.8).pdf
Files
(559.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:a6a492c6fbb963ede58a2157ee0e9a0d
|
559.0 kB | Preview Download |
Additional details
Dates
- Created
-
2026-05Revised draft
Software
- Repository URL
- https://github.com/humanaiconvention/prism/tree/main/experiments/sgt
- Programming language
- Python
- Development Status
- Active
References
- Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., & Baraniuk, R. G. (2023). Self-Consuming Generative Models Go MAD. arXiv preprint arXiv:2307.01850. https://arxiv.org/abs/2307.01850
- Dohmatob, E., Feng, Y., Yang, P., Charton, F., & Kempe, J. (2024). A Tale of Tails: Model Collapse as a Change of Scaling Laws. Proceedings of the 41st International Conference on Machine Learning (ICML). https://proceedings.mlr.press/v235/dohmatob24b.html
- Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., Sleight, H., Hughes, J., Korbak, T., Agrawal, R., Pai, D., Gromov, A., Roberts, D. A., Yang, D., Donoho, D. L., & Koyejo, S. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv preprint arXiv:2404.01413. https://arxiv.org/abs/2404.01413
- Feng, Y., Dohmatob, E., Yang, P., Charton, F., & Kempe, J. (2024). Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification. arXiv:2406.07515. https://arxiv.org/abs/2406.07515
- Keisha, F., Wu, Z., Wang, Z., Koshiyama, A., & Treleaven, P. (2025). Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training. arXiv:2509.04796. https://arxiv.org/abs/2509.04796
- Rajaee, S., Pratik, K., Cesa, G., & Behboodi, A. (2025). Local Look-Ahead Guidance via Verifier-in-the-Loop for Automated Theorem Proving. arXiv:2503.09730. https://arxiv.org/abs/2503.09730
- Schaeffer, R., Kazdan, J., Arulandu, A. C., & Koyejo, S. (2025). Position: Model Collapse Does Not Mean What You Think. arXiv:2503.03150. https://arxiv.org/abs/2503.03150
- Shi, L., et al. (2025). A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective. arXiv:2509.16499. https://arxiv.org/abs/2509.16499
- Shukla, A., Knowles, S., Madugula, M., Farris, D., Angilly, R., Pombo, S., Xu, A., An, L., Balasubramanian, A., Yu, T., Ren, J., & Akkiraju, R. (2025). Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement. arXiv:2510.27051. https://arxiv.org/abs/2510.27051
- Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755-759. https://doi.org/10.1038/s41586-024-07566-y
- Straňák, P. (2026). Entropic Limits of Iterative Computation in Generative AI: Model Collapse Explained by the Data Processing Inequality and the AI Theorem. Preprints.org, version 2. https://doi.org/10.20944/preprints202507.2260.v2