Published July 10, 2026
| Version v0.5.4
Preprint
Open
The Width of a Voice: Placing Machine Imitation Inside Authors' Own Variation
Description
Whether a language model can write as a specific author is usually put to a detector: generate an imitation, ask whether it passes. That frame yields verdicts, not measurements; it cannot say how close an imitation is, or compared to what. The missing comparator is the one stylometry uses for humans: how much the author varies from herself, at the length measured. An imitation success rate against an author of unstated width is a numerator without a denominator. We build the denominator: a measurement space calibrated on 15 published novelists under Burrows's Delta, with per-author, length-matched envelopes of within-author variation, validated by a positive control — the authors' own held-out text re-enters its own envelope at an observed 84–88%, the yardstick every entry rate is read against. Four findings. Prompting a model with an author's name roughly triples entry into the author's envelope, measured on the closed-class function words that classical attribution rests on (10.7% to 30.5%); a cross-target control finds the increment substantially target-specific (text styled for one author enters other authors' envelopes with no detectable difference from never-prompted text). What style prompting transfers is mostly surface texture — the underlying function-word machinery barely moves in aggregate. Giving a model the author's own opening, with no author named, beats naming the author for every informative model that complied; the best style imitator refused every such request, leaving the field's strongest condition partly unmeasurable on frontier models. Imitability is substantially a property of the author: styled text lands at nearly the same distance from whichever target it aims at, so envelope width — how much an author varies from herself — predicts what her envelope admits (r = +0.84). Popular "AI tells" (em dashes, triads, hedging) barely separate machine fiction from celebrated novelists: a combined twelve-tell detector is a coin flip (AUC 0.506) and, tuned to catch half the machine samples, flags 50.8% of the human windows — a caution for anyone applying checklists to people. We release the instrument, corpus (including refusals), envelopes, and a public-domain replication package.
Notes
Files
width_of_a_voice_preprint_v0.5.4.pdf
Files
(938.2 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:47a0461e26446f636de89dc3c7b2f7ed
|
938.2 kB | Preview Download |
Additional details
Related works
- Is supplemented by
- Software: https://github.com/wbryanta/author-manifold (URL)
- Software: https://github.com/wbryanta/mnemosyne-overview (URL)
Dates
- Updated
-
2026-07-100.5.3 to 0.5.4 -- cross-target matrix, LaTeX edition, Figure F8, new citations.