Published August 29, 2026 | Version 1.0

The Personalization Gap: How a Model's Knowledge of the User Reshapes Brand Recommendations in Generative AI

  • 1. Cassie Clark Marketing
  • 2. friction AI

Description

Generative answer engines increasingly mediate brand discovery: a user asks for "the best X," and a large language model (LLM) returns a short list of named brands. A scalable way to measure this surface is an application-programming-interface (API) call without an account or persistent user history. Real users, by contrast, often reach these systems through logged-in sessions that may carry account history. This paper reports a controlled study of how much a model's knowledge of the user changes the brands it recommends, and what history-free API measurement may therefore miss. Across four LLMs (ChatGPT, Claude, Gemini, Perplexity), six personas in three consumer categories, and thirty prompts, the same questions were asked in three ways: through a web-enabled API with no user account or history; through a logged-in temporary session that did not retain history (Temp); and through a logged-in session with a controlled usage history (history-primed). All three were configured for the UK market and run from UK IP addresses. The study comprises 8,609 successful collections over nine days. This paper terms the central effect the Personalization Gap: a statistically detectable change in the reliably recommended brand shortlist that is associated with user context and is partly missed by history-free API measurement. Most recurring brands stayed the same: the average history-primed core contained about 3.7 recurring brands, of which roughly 2.8 were shared with the neutral session and 0.9 differed. After allowing for normal run-to-run variation through a within-model control called the Volatility Floor, the history-primed brand set diverges significantly for ChatGPT (+16.0 percentage points, p < 0.001) and Gemini (+33.3pp, p < 0.001); Perplexity is borderline on this primary comparison (+7.2pp, one-sided p = 0.034), while Claude shows no detected brand-set difference (+1.0pp, p = 0.412) and instead changes the search queries it uses.

The clearest single dimension is geographic: the history-primed users in this UK study received a greater UK/EU-origin and smaller United States (US)-origin brand share (+6.6pp / −4.9pp) and more .uk-domain citations (+6.0pp). Across 173 matched comparisons evaluated at an equal depth of up to sixteen repeated responses per series, cold-API measurement recovers about half of the history-primed core from one response and reaches 79% recall at the endpoint, whereas Temp reaches 95%. Whether that observed gap persists at greater within-series depth is stated as a testable Accuracy Ceiling hypothesis, not as an established asymptote.

The paper closes with a tier-and-confidence rule for reporting brand visibility on a surface that varies from run to run. The deposit includes the paper, the as-executed protocol, aggregate evidence for every published figure and table, a claim-level provenance index, and checksums. Response-level outputs, unpublished analyses, code, and proprietary measurement materials remain confidential.

Files

personalization-gap-v1.0.pdf

Files (916.0 kB)

Name Size Download all
md5:0cb4c210f4d1cf39ee33ad892cd562bf
27.2 kB Preview Download
md5:0dcdc2acc177504ef3b64879b3faffaa
4.2 kB Preview Download
md5:8a8ca1e1c4fa9828a8199711f35ff9b1
422 Bytes Preview Download
md5:5d11bfebdb80fdf5c92993b75a90d59c
868.9 kB Preview Download
md5:62a139689dd15503c04803d37348b7b3
13.4 kB Preview Download
md5:0964d0da6e5fd307cf5d24bfa64ddc7f
1.3 kB Preview Download
md5:4b87f0b7f5a7216ae01d28403fbff350
531 Bytes Download