Published August 27, 2026
| Version 7
Working paper
Open
An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026
Description
Excess vocabulary, a word's frequency above its pre-2023 trend, has been used to measure how language models changed scholarly English. We adapt it to Korean with morphological units on 398,296 KCI abstracts (2018 to August 2026), and 47,165 Vietnamese abstracts as an exploratory comparison. Placebo floors are 0.1–2.2 points for the single-word statistic and at most 2.9 for the re-selected split-half set statistic. Korean abstracts show nothing in 2023, onset in late 2024, a rise through 2025 flattening in mid-2026: 시사하다 "suggest" appears in 21.4% of 2026 abstracts against 5.3% expected; plain verbs like 알아보다 "look into" fall to a quarter of trend. Under stated identification assumptions the single-word conditional lower bound on LLM-processed abstracts is 3.5% in 2024, 10.5% in 2025 and 16.1% in 2026; a split-half set bound is 7.8%, 20.6% and 33.0%. Tested subject-matter controls do not explain it: restricting the set to lemmas three language-model annotators all call style leaves 14.7 points, and pairing each 2026 abstract with its journal's closest base-period abstract leaves the difference at 34.1. Tested translation routes do not explain it: the surface marks of translated Korean fall as the markers rise. In the same articles' English abstracts the excess appears a year earlier; where the English side carries no markers the Korean shift persists at 30 to 66% of its uncorrected rate. Control abstracts from three providers reproduce the rising words and show generation-associated marker turnover; implied prevalences are scenario-dependent.
Working paper, version 7 (27 August 2026). Version 7 states the identification assumptions behind the conditional lower bound, reports implied prevalences as prompt- and model-conditioned scenarios, adds an independent style/topic annotation of the marker lemmas with restricted bounds (Section 5.13), a topic-matched within-journal comparison, a translation-route check with model translation conditions (Section 5.14), the measured sensitivity of the English marker set (Section 5.12), a longitudinal journal-cluster bootstrap with stated replicate counts, a 100-split cross-fitting check (Appendix F), and a clean-room reproducibility package that rebuilds every table, figure and the manuscript from the shipped results; the whole analysis is computed on the completed harvest of 398,296 abstracts. The version history is in Appendix J. The reproducibility package contains the analysis code, per-year document-frequency tables for Korean, English and Vietnamese, the generated control abstracts (1,437 Korean in sixteen conditions and 514 English in six, from OpenAI, Anthropic and EXAONE models), the annotation, matching, translation-route and robustness results, the results manifest and the manuscript consistency gate. A public Korean AI-style dictionary and checker built from the same measurements are at https://os.intframe.com/report/ai-style-dictionary-ko and https://os.intframe.com/report/ai-style-check-ko.
Notes
Files
ai-style-lexicon-reproducibility-v7.zip
Files
(52.7 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:4412935597bc769ddad14f40bd06beac
|
51.1 MB | Preview Download |
|
md5:6a9743aab9d48aee70d9fcaf6415a733
|
1.5 MB | Preview Download |