Published June 24, 2026 | Version v1

How LLMs Source Brand Reputation Across Languages and Markets: A Cross-Market Citation Dataset (2026)

  • 1. Rankfor.AI
  • 2. ROR icon Estonian Entrepreneurship University of Applied Sciences

Description

The citations that grounded large-language-model answers about brands, merged from three Rankfor.AI studies, supporting the paper "How Large Language Models Source Brand Reputation Across Languages and Markets." It covers 167,551 URL-grounded citations (189,974 total attribution rows) across 128 brands, 13 languages, and 12 home markets, from three grounded models (GPT, Gemini, Perplexity). Each citation carries its domain, source type, language, and model, with the citation-title field that resolves Google grounding redirectors to the real publisher. Headline results: AI grounds brand answers in third-party sources 85.7% of the time and in the brand's own site 14.3%; 80% of citations come from about 18% of domains (Zipf alpha 0.86, R^2 0.983); Wikipedia is the most-cited domain in 11 of 12 languages, with Lithuanian (vz.lt) the exception; in Poland the top domain is YouTube and four HR/careers portals out-cite Polish Wikipedia about 2 to 1; Perplexity is the highest-volume citer. Includes the analysis ledger and a reproduction script.

Files

llm-sourcing-cross-market-2026-zenodo.zip

Files (10.6 MB)

Name Size Download all
md5:db8ba8b3c52df068e5e845b5f1664bb0
10.6 MB Preview Download

Additional details

References

  • arXiv:2606.23165 (The Language Blind Spot; the answer-text companion built on the same Nordic-Baltic data)
  • arXiv:2606.23057 (Who Owns the AI Recommendation?; the recommendation-share study this work extends)