The SEO Floor: Measuring Google Rank Distribution of AI-Cited Pages
Authors/Creators
Description
Whether Generative Engine Optimization (GEO) is a discipline distinct from Search Engine Optimization (SEO), or merely SEO repackaged, has been debated since the rise of consumer-facing AI chat platforms. Empirical resolution requires measuring two things: where on Google's search results AI-cited pages actually rank, and whether content features pre-registered as "GEO levers" predict citation independently of Google rank. We collected 100,411 AI citation events from four production AI platforms (ChatGPT, Perplexity, Claude, Google AI Mode) across 2,000 user queries, and assembled a comparison pool of 165,661 unique URLs from Google top-100 SERPs for the same queries. Combining the citation corpus with the comparison pool yields 114,729 (URL, query) observations for a mixed-effects logistic regression of citation probability on Google rank tier and GEO composite score, with query random intercepts. **Headline finding 1**: While 75.4% of citation events are aggregated to pages outside Google's top 30, *per-page* citation odds span a 34× range across rank tiers — a top-3 page is **7.82× more likely to be cited (OR vs Tier 3, 95% CI 7.28–8.39)** while a rank 31–100 page is **4× less likely (OR=0.23, 95% CI 0.22–0.24)**. The aggregate "75% deep-tier" figure is a denominator artifact: the internet beyond rank 30 has vastly more pages than the top 30, so even at low per-page citation probability the absolute count of deep-tier citations dominates. The two figures must be reported together. **Headline finding 2**: A pre-registered GEO composite of seven content features adds small but real predictive power above Google rank (Z-sum OR=1.06; PCA-1 OR=1.15 per 1 SD), driven primarily by schema markup (OR=1.31). Statistics density and list structure show small positive effects after methodological correction; heading density shows a small negative effect. **Headline finding 3**: The 75% Tier-4+ aggregate is overwhelmingly composed of URLs Google ranks beyond #100 (90% of "Tier 4" events) — not the 31–100 band that the H2 regression speaks to — and these deep-tier citations are 77% one-hit-wonders cited by a single AI platform, with sharp platform divergence on user-generated-content tolerance (Claude 0.6% UGC in deep tier; Perplexity 24%). The Lily-Ray-aligned framing ("AI citation is gated by Google ranking") is empirically supported, with schema markup emerging as the strongest single content-feature predictor inside the gate; whether AI parsers consume schema directly or schema is a proxy for site quality is unresolved by observational data and is the target of a planned follow-up interventional study.
Files
study_a_paper_draft.pdf
Files
(340.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:f360fac1328be3cd09889c41d8d629c2
|
340.5 kB | Preview Download |
Additional details
Related works
- Is documented by
- Other: 10.17605/OSF.IO/FMSRD (DOI)
- Is supplemented by
- Dataset: 10.5281/zenodo.19787328 (DOI)