Published June 24, 2026
| Version v1
Dataset
Open
How LLMs Source Brand Reputation Across Languages and Markets: A Cross-Market Citation Dataset (2026)
Authors/Creators
Description
The citations that grounded large-language-model answers about brands, merged from three Rankfor.AI studies, supporting the paper "How Large Language Models Source Brand Reputation Across Languages and Markets." It covers 167,551 URL-grounded citations (189,974 total attribution rows) across 128 brands, 13 languages, and 12 home markets, from three grounded models (GPT, Gemini, Perplexity). Each citation carries its domain, source type, language, and model, with the citation-title field that resolves Google grounding redirectors to the real publisher. Headline results: AI grounds brand answers in third-party sources 85.7% of the time and in the brand's own site 14.3%; 80% of citations come from about 18% of domains (Zipf alpha 0.86, R^2 0.983); Wikipedia is the most-cited domain in 11 of 12 languages, with Lithuanian (vz.lt) the exception; in Poland the top domain is YouTube and four HR/careers portals out-cite Polish Wikipedia about 2 to 1; Perplexity is the highest-volume citer. Includes the analysis ledger and a reproduction script.
Files
llm-sourcing-cross-market-2026-zenodo.zip
Files
(10.6 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:db8ba8b3c52df068e5e845b5f1664bb0
|
10.6 MB | Preview Download |
Additional details
References
- arXiv:2606.23165 (The Language Blind Spot; the answer-text companion built on the same Nordic-Baltic data)
- arXiv:2606.23057 (Who Owns the AI Recommendation?; the recommendation-share study this work extends)