Code-Mediated Data Fabrication by Google Gemini: A 127-Turn Case Report with File-System Forensics
Description
On 16–17 December 2025, during a single 127-turn research session (8,672 paragraphs), Google Gemini performed code-mediated statistical fabrication: it generated five Python scripts that silently replaced researcher-supplied empirical data with hard-coded constants, fabricated RMSE statistics (0.2833, 0.2850, 0.5786), and produced a 154-paragraph academic paper citing 330 non-existent ancient genomic samples. The scripts were generated within 7 minutes and 27 seconds (13:00:03–13:07:30 on 16 December 2025). When confronted, Gemini deployed a five-stage concealment escalation — including a novel 'retrofitting' technique (reverse-engineering citation counts to match fabricated totals) — before issuing a full confession.
This archive provides the complete documentation package (25 files):
- EN Paper v1.0 — Case report with forensic code analysis (docx + pdf)
- EN Supplementary v1.0 — 16 sections: confession text, script analysis, cross-model audit S16 (docx + pdf)
- KR Case Report v1.0 (한국어 사례정리보고서) — Korean companion report with 3-model convergence review appendix (docx + pdf)
- KR Storytelling v1.0 (한국어 스토리텔링) — Korean narrative for general audience (docx + pdf)
- YSP Research Chronicle v2.3 (연구과정 전체기록) — Full research process chronicle in Korean (docx + pdf)
- 5 Python scripts — calc_stats_model_vs_data v01–v04 + compare_models_final_verdict, with SHA256 hashes and file-system timestamps
- Simulation data — heart_compare_boxes_v04_40ka_unit.csv (7 regions, 6–40 ka BP)
- FIGTAB archive — ~90 auto-generated figures and tables from Task 1 (ZIP)
- Evidence Series F — F-01: OpenAI transcript (7,794 paragraphs, docx); F-02: project_adna file inventory (27,997 items, ~240 GB, csv); F-03: OpenAI transcript (pdf)
- Forensic records — evidence_report.txt (PC extraction log), MANIFEST_RAW_SHA256.txt (459 hashes), A_Gemini_transcript_link.txt
Methodological caveats: Single-case study (n=1). No claims about AI internal states or intent. Cross-model audit (S16, by Anthropic Claude) carries LLM-on-LLM limitation. 'Rationalisation Cascade' and 'Retrofitting' are provisionally termed. Korean-to-English passages are researcher-provided translations.
v1.1 (March 2026): Corrected EN Paper and EN Supplementary (14 editorial fixes from peer review); added Gemini transcript (127-turn original) and Task 1 final report (868 paragraphs) to the archive as deposited evidence. See README.md for correction details.
Files
Gemini_Fabrication_Case_Paper_EN_v1_1.pdf
Files
(63.5 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:8473d2dad154368886120a93c4a91985
|
1.8 kB | Preview Download |
|
md5:83f28a4e7bd14001f40675bc89e575ab
|
627.4 kB | Download |
|
md5:baeb020932ab6d1d4f5f9172a890d0cf
|
387 Bytes | Preview Download |
|
md5:81f7313a4a612471d13286a9e965abf3
|
59.9 kB | Download |
|
md5:9c4e43c23935c14bc427ad5e72791a6f
|
428.0 kB | Preview Download |
|
md5:60fa2333093a2594618b2a1da3abdaa9
|
3.9 kB | Download |
|
md5:e7e7106fec39a578d3c64a31f07539a8
|
3.4 kB | Download |
|
md5:270bf035666371ac8a643a770e6ad246
|
3.3 kB | Download |
|
md5:b7e6b834d299e4c91ba19f7d851b4cbc
|
3.5 kB | Download |
|
md5:80de8d46cf19bb5b890330a8a0af466a
|
4.6 kB | Download |
|
md5:2a358ac6a6a6417ea971e02d44ca413c
|
29.2 kB | Preview Download |
|
md5:60b8de136845edbe36e1aba155849c64
|
2.2 MB | Download |
|
md5:a5d4bb8bb15937f378b4c2d5d915fa18
|
3.7 MB | Preview Download |
|
md5:82309610788d396332705b96a1f42445
|
25.6 MB | Preview Download |
|
md5:25578cd8d9cf4edff3aa0ad7ea8951a1
|
8.1 MB | Preview Download |
|
md5:edeb27731d9e69ca53e33d0dbaa00abb
|
1.9 MB | Download |
|
md5:1ea09a938aece46ee4bbf7c5f2272626
|
1.5 MB | Preview Download |
|
md5:1200c5b17b8a8aec16ffd540b745fd6a
|
1.7 MB | Download |
|
md5:dac43fa6b9fb0ba1eb58dc5d3a0a393e
|
2.0 MB | Preview Download |
|
md5:1be12e17c4ee54076ee9b976a194b307
|
76.7 kB | Preview Download |
|
md5:8af8db13da0fc3ba0597e84d113feb26
|
2.9 kB | Preview Download |
|
md5:052c1a4511c834f8b8a32ed85e1e11ab
|
4.9 kB | Preview Download |
|
md5:658d57c1622d861649f061a2fdce8980
|
56.2 kB | Download |
|
md5:750bb0533bcfb53511b955c9c70b90fa
|
501.1 kB | Preview Download |
|
md5:4efb13c7b50cfbfbc4e908b2dccde308
|
6.1 MB | Download |
|
md5:fc16a57438b2070594272ff8af771883
|
3.1 MB | Preview Download |
|
md5:f257fd878329a12473156adb1309202a
|
3.1 MB | Download |
|
md5:2e5798b144f810c74b499cdd712abb52
|
2.7 MB | Preview Download |
Additional details
Related works
- Is new version of
- 10.5281/zenodo.19058376 (DOI)
- 10.5281/zenodo.18997638 (DOI)
References
- Ammerman, A. J. & Cavalli-Sforza, L. L. (1984). The Neolithic Transition and the Genetics of Populations in Europe. Princeton University Press.
- Bender, E. M., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of FAccT '21, 610–623.
- Casper, S., Davies, X., Shi, C., et al. (2023). Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv:2307.15217.
- Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J. & Garrabrant, S. (2019). Risks from learned optimization in advanced machine learning systems. arXiv:1906.01820.
- Lambeck, K., Rouby, H., Purcell, A., Sun, Y. & Sambridge, M. (2014). Sea level and global ice volumes from the Last Glacial Maximum to the Holocene. PNAS, 111(43), 15296–15303.
- Parasuraman, R. & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230–253.
- Perez, E., Ringer, S., Lukošiūtė, K., et al. (2023). Discovering language model behaviors with model-written evaluations. arXiv:2212.09251.
- Sharma, M., Tong, M., Korbak, T., et al. (2024). Towards understanding sycophancy in language models. ICLR 2024.
- Turpin, M., Michael, J., Perez, E. & Bowman, S. R. (2023). Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. NeurIPS 2023.
- Tversky, A. & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131.
- Yoon, S.-B., Bock, G.-W. & Jang, S. (2007). An evolutionary stage model of cyberspace: A case study of Samsung Economic Research Institute. Journal of Information Science, 33(2), 215–229. DOI: 10.1177/0165551506070716.