Published August 29, 2026
| Version 1.0
Report
Open
The Personalization Gap: How a Model's Knowledge of the User Reshapes Brand Recommendations in Generative AI
Authors/Creators
- 1. Cassie Clark Marketing
- 2. friction AI
Description
Generative answer engines increasingly mediate brand discovery: a user asks for "the best X," and a large language model (LLM) returns a short list of named brands. A scalable way to measure this surface is an application-programming-interface (API) call without an account or persistent user history. Real users, by contrast, often reach these systems through logged-in sessions that may carry account history. This paper reports a controlled study of how much a model's knowledge of the user changes the brands it recommends, and what history-free API measurement may therefore miss. Across four LLMs (ChatGPT, Claude, Gemini, Perplexity), six personas in three consumer categories, and thirty prompts, the same questions were asked in three ways: through a web-enabled API with no user account or history; through a logged-in temporary session that did not retain history (Temp); and through a logged-in session with a controlled usage history (history-primed). All three were configured for the UK market and run from UK IP addresses. The study comprises 8,609 successful collections over nine days. This paper terms the central effect the Personalization Gap: a statistically detectable change in the reliably recommended brand shortlist that is associated with user context and is partly missed by history-free API measurement. Most recurring brands stayed the same: the average history-primed core contained about 3.7 recurring brands, of which roughly 2.8 were shared with the neutral session and 0.9 differed. After allowing for normal run-to-run variation through a within-model control called the Volatility Floor, the history-primed brand set diverges significantly for ChatGPT (+16.0 percentage points, p < 0.001) and Gemini (+33.3pp, p < 0.001); Perplexity is borderline on this primary comparison (+7.2pp, one-sided p = 0.034), while Claude shows no detected brand-set difference (+1.0pp, p = 0.412) and instead changes the search queries it uses.
The clearest single dimension is geographic: the history-primed users in this UK study received a greater UK/EU-origin and smaller United States (US)-origin brand share (+6.6pp / −4.9pp) and more .uk-domain citations (+6.0pp). Across 173 matched comparisons evaluated at an equal depth of up to sixteen repeated responses per series, cold-API measurement recovers about half of the history-primed core from one response and reaches 79% recall at the endpoint, whereas Temp reaches 95%. Whether that observed gap persists at greater within-series depth is stated as a testable Accuracy Ceiling hypothesis, not as an established asymptote.
The paper closes with a tier-and-confidence rule for reporting brand visibility on a surface that varies from run to run. The deposit includes the paper, the as-executed protocol, aggregate evidence for every published figure and table, a claim-level provenance index, and checksums. Response-level outputs, unpublished analyses, code, and proprietary measurement materials remain confidential.
Files
personalization-gap-v1.0.pdf
Files
(916.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:0cb4c210f4d1cf39ee33ad892cd562bf
|
27.2 kB | Preview Download |
|
md5:0dcdc2acc177504ef3b64879b3faffaa
|
4.2 kB | Preview Download |
|
md5:8a8ca1e1c4fa9828a8199711f35ff9b1
|
422 Bytes | Preview Download |
|
md5:5d11bfebdb80fdf5c92993b75a90d59c
|
868.9 kB | Preview Download |
|
md5:62a139689dd15503c04803d37348b7b3
|
13.4 kB | Preview Download |
|
md5:0964d0da6e5fd307cf5d24bfa64ddc7f
|
1.3 kB | Preview Download |
|
md5:4b87f0b7f5a7216ae01d28403fbff350
|
531 Bytes | Download |