AI Recommendation Calibration Study 2026: aggregate results for UK local business recommendations across four large language models
Authors/Creators
Description
What this dataset is. Aggregate results from a controlled calibration study measuring how four large language model (LLM) products recommend local UK businesses when asked category-level questions such as "best dentist in Bristol" or "emergency plumber Manchester". Between May 2026 and July 2026, 4,800 model responses were collected across a fully crossed factorial design: two business categories × three UK cities × four AI platforms × four location-phrasing modes × five prompt intents per category × ten repeats per combination.
Headline findings.
- 3,583 distinct businesses were recommended across the study.
- 64.1% of distinct businesses appeared in only a single response.
- Cross-platform Jaccard similarity of recommendation sets remained below 0.15 across all tested cells, indicating that AI-generated recommendations are highly fragmented rather than converging on a common shortlist.
- A cross-model reviewer agreement check on 100 random runs (Claude Haiku 4.5 vs. GPT-4o-mini) produced Jaccard 0.61 on entity extraction and 0.63 on signal classification, with per-category Cohen's kappa reported in the methodology PDF.
Platforms tested. OpenAI GPT-5.3 Chat, Anthropic Claude Sonnet 4.6, Google Gemini 3 Flash, Perplexity Sonar (search-enabled). All accessed via commercial API endpoints with provider pinning and fallbacks disabled.
Files in this deposit.
summary_stats.json— study manifest with counts, categories, cities, platforms, headline numbers and aggregation rules.recommendations_by_category.csv— response volume, distinct businesses recommended, average recommendations per response, positive-sentiment share, per category × city × platform.concentration_by_category.csv— top-1 share, top-3 share, HHI, share of businesses recommended only once, per category × city × platform.cross_platform_overlap.csv— pairwise Jaccard and overlap coefficient across platforms, per category × city.signal_frequency.csv— frequency of each reasoning signal across a 16-category taxonomy, per category × platform.methodology.pdf— full methodology, study design, statistical measures, cross-model reviewer agreement, and limitations.
Privacy. All published files are aggregates. No business names, competitor names, cited sources, verbatim AI response text, postcodes, neighbourhoods, or evidence quotes are included. Cells backed by fewer than five underlying observations are suppressed to reduce the risk of indirectly identifying individual businesses.
Suggested citation. Billingham, I. (2026). AI Recommendation Calibration Study 2026: aggregate results for UK local business recommendations across four large language models [Data set]. Zenodo.
Contact. hello@ai-mention.co.uk
Files
concentration_by_category.csv
Files
(43.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:5196e8a32da16544b25da737bd3d2b25
|
2.1 kB | Preview Download |
|
md5:41b6c27de30af7cfa482654790110f17
|
3.1 kB | Preview Download |
|
md5:b2900f343782a8443dc2785b95f5a881
|
27.2 kB | Preview Download |
|
md5:8aac35832218d000047161a2bc26046f
|
2.0 kB | Preview Download |
|
md5:8557b57154bb176f1c0fd256cfa859d2
|
7.9 kB | Preview Download |
|
md5:f343b3a14501731109164ab5005582cd
|
1.0 kB | Preview Download |
Additional details
Related works
- Is documented by
- Other: https://ai-mention.co.uk (URL)