We analyzed 21,571 URL citations from six AI models and found eight things. A citation is every link a model shows as a source for its answer: if one answer lists seven sources, we count seven. The most uncomfortable of the eight: no brand site cleared 10% of the citations in its own category. The most useful one: the typical citation is not your homepage, it is a specific internal page. And the one that separates us most from what is already published: blog content is the most cited type by volume and second to last by position inside the source list. This article walks through the eight findings with their strongest data point, and each one links out to the full analysis.
All 20 pages, with every table and the 26 references.
The typical citation is not your homepage
01. Models cite internal pages that are clean and specific
Between 67.8% and 75.2% of citations point to pages sitting two or more levels deep inside the site. Not
example.com and not example.com/blog, but
example.com/blog/how-to-measure-geo. 99.98% of citations use HTTPS, 94.1% do not end in a file
extension, and only 4.2% carry parameters in the address.
This separates two things people tend to conflate: a model citing your company, or a model citing a specific answer you published. The second is far more common, and that changes where the investment belongs.
What makes the finding strong is not the percentage, it is the repetition. The spike at depth 2 shows up just as sharply in both datasets, and those two datasets share only 3.7% of their URLs. It is not an artifact of one instrument or one category. The implication is boring and hard to sell: a good chunk of AI readiness work is architecture hygiene. Clean, stable, specific addresses. There is no trick here, there is an entry requirement plenty of sites still do not meet.
Read the full analysis in the study, section 02
02. Blogs win volume and lose position
When a model answers, it lists its sources in an order, and that order is not cosmetic: the first one is the most visible and, in several interfaces, the only one you see without expanding the rest. Almost no published study measures that. They measure whether you showed up. We added where, and the picture changes.
| Page type | Citations | Reaches 1st slot | Average position |
|---|---|---|---|
| Homepage | 1,165 | 14.2% | 5.9 |
| Guide or resource | 2,436 | 12.5% | 8.4 |
| Social and video | 1,835 | 11.4% | 7.5 |
| Product page | 4,841 | 8.6% | 8.5 |
| Company blog | 8,139 | 7.5% | 10.2 |
| 411 | 6.8% | 7.4 |
The company blog has twenty times the volume of reference sites and less than a third of their first slot rate. It is the most cited content type in the study and among the last in quality of placement. And placement matters more than it looks, because clicks are vanishingly rare: Pew Research found that only 1% of users click any link inside an AI Overview. If nobody clicks, the value of a citation concentrates in how visible its slot is.
Blogs dominate AI citations and blogs win AI citations are two different claims, and the industry uses them as synonyms. We propose splitting reach (do we show up?) from quality of placement (do we show up first?). Blogs optimize the former and are mediocre at the latter. This does not say stop blogging. It says: stop measuring mentions alone, and do not expect the blog to hand you the top slot, because structurally it does not.
Read the full analysis in the study, section 03
Four models mostly cite the blog. ChatGPT does not
03. There is no such thing as one URL for every LLM
Any combination of models we take agrees on barely 3.5% of the exact URLs it cites. At the domain level the overlap rises to 10.4%. Put another way: they may trust the same sites, but they rarely pick the same page.
| Model | Citations | Blog | Product | Homepage | Social and video |
|---|---|---|---|---|---|
| Claude | 2,909 | 42.5% | 28.2% | 6.9% | 0.5% |
| Google AI Mode | 2,514 | 39.1% | 23.4% | 5.6% | 12.1% |
| ChatGPT | 2,151 | 12.6% | 31.3% | 10.6% | 0.7% |
| Gemini | 1,567 | 38.7% | 22.4% | 16.5% | 3.1% |
| Perplexity | 1,477 | 33.1% | 30.6% | 7.0% | 4.7% |
Classification by page type and site type over the 10,618 citations of the multi model set. 97% of the citations were identified.
Four of the five mostly cite the company blog, at weights ranging from 33.1% to 42.5%. ChatGPT does not: its most cited category is product pages, at 31.3%, while its blog reaches only 12.6%. Add guides, PDFs and documentation and 62% of what ChatGPT cites is product material.
The other wide gap is social, and we now know what it is made of: Google AI Mode runs at 12.1%, split into 6.8% social networks, 3.6% video and 1.7% LinkedIn. Claude adds up to 0.5% across all three. A twenty-fourfold difference. That is not a percentage tweak, it is a different content priority.
The commercial consequence is finer than it sounds: you cannot promise the same URL wins on every LLM. There is shared ground, especially on product pages. What you build is a portfolio: types that serve several models, formats where you specialize, and measure model by model.
What each model cites, one by one
04. Nobody owns the conversation, and your site least of all
The single most cited domain in the entire study accounts for 3.6% of volume. Not even YouTube, which leads it, reaches 4%. You need 500 domains to cover half the citations. There is no oligopoly: the long tail is the market, with 7,282 domains competing.
Now the number that matters. When we measure how many citations in a category point at the site of a brand in that same category, no brand domain cleared 10%, and the average was 3.3%. The remaining 96.7% goes to third parties: media, forums, institutions, platforms and competitors. Your brand is not in the conversation. It is being mentioned by others, or it is not there.
The industry reports that brand sites account for between 38% and 70% of citations, depending on the model. That number and ours are both correct, and they measure different things. Brand sites is simply the sum of every commercial site in the world: a company, its competitors, and any business relevant to the question. That half of ChatGPT citations go to corporate sites says nothing about how many go to yours. If someone hears 50% and their dashboard shows 4%, the natural conclusion, and the wrong one, is that their team is failing.
The work of Chen, Wang, Chen and Koudas, at the University of Toronto, documents the same thing from another angle: a systematic bias in AI search engines toward authoritative third party sources, over brand owned content. The practical consequence is not to give up. It is to stop treating your own site as the primary AI visibility channel and start treating it as one among several, alongside presence in media, in video, and in the institutional sources of the category.
Read the full analysis in the study, section 05
With the methodology, the tables by industry, and what the study does not say.
The two faces of Google
Google has two surfaces that behave very differently and that were being measured as if they were one: the AI Overview, the summary that shows up above the results as soon as you search, and AI Mode, the follow-up chat you land in when you ask again. The three findings that follow come from a parallel experiment: 518 Google searches, between July and August 2026, covering ten industries and two markets. For each one we captured all four surfaces from the same screen and the same session, using GEO Observer: the AI Overview, the organic results, the paid ads, and AI Mode on the follow-up. That design is what makes it possible to compute overlap per page within the same search, instead of comparing captures taken at different moments.
05. AI Overview repeats your SEO. AI Mode throws it out
47.45% of AI Overview citations were already in the organic results of that same search. Ask a follow-up and move to AI Mode, and only 19.98% of those organic pages survive. And AI Mode does not just discard: it also multiplies. The average goes from 7.53 citations per AI Overview to 14.98 per AI Mode turn, and 85.29% of those citations were not in the organic results, the ads, or the AI Overview of that search.
This resolves a contradiction the industry has been carrying for two years. When a study reports that 76% or 97% of AI citations come from the organic top 10, it is looking at AI Overviews. When another reports 12% or 17%, it is looking at AI Mode or ChatGPT. They do not contradict each other: they describe different surfaces of the same system, and the label AI citations throws them all in the same bucket.
Read the full analysis of the 518 searches
06. Paid ads do not exist for the generative answer
Ads accounted for 0.03% of AI Overview citations. And persistence into AI Mode was 0% across 75 paid pages, without a single exception in ten industries. The most plausible explanation is one of genre, not platform policy: an ad is written to convert someone who already decided to click, and a citation has to read like a source that answers the question.
The budget implication is concrete. If someone justifies ad spend with the argument of presence in generative answers, this data contradicts them directly.
Read the full analysis of the 518 searches
07. AI Overview overweights video far beyond SEO
YouTube weighs 1.72% in organic results and 7.72% in AI Overview citations. More than four times. Add social media and it reaches 16.08%. Reddit and Wikipedia together, in the same sample, add up to 1.57%. There is a received idea that generative answers favor forums and encyclopedias, and in this data it is the reverse by an order of magnitude.
AI Overview does not inherit the proportions of SEO. It actively overweights video and social content. If your category lends itself to that, a well structured video can weigh more than an article optimized the old way.
Read the full analysis of the 518 searches
AEO has a floor, not a staircase
08. There is a threshold to cross, and then a noisy plateau
An AEO score is a 0 to 100 grade for how ready a page is to be cited. It combines technical, formatting and identity signals. Almost every tool in the space, ours included, computes something similar, and the implicit promise is that raising the score raises the odds of being cited. We crossed 1,486 URLs with known scores against the citations they actually received, and that promise holds only in one stretch.
| Page grade | Citation rate |
|---|---|
| A · 85 and up | 11.9% |
| B · 70 to 84 | 2.4% |
| C · 55 to 69 | 17.5% |
| D · 40 to 54 | 2.8% |
| F · under 40 | 0% |
The floor is absolute: no page below 40 points received a single citation. Above it, the score stops sorting anything. Pages graded C get cited seven times more often than pages graded B, which score better. There is no staircase that climbs with the score. There is one step and then a noisy plateau.
The most uncomfortable data point sits in the formatting component. Ranking articles by their structure and presentation score, the relationship with real citation runs backwards: the worst scoring quartile gets cited 9.1% of the time, and the best scoring quartile, 0%. The likely explanation is not that formatting hurts, but that the score captures something else. Pages that score 100 on structure are long, exhaustive articles, the kind a brand produces to rank, and what a model extracts is a short fragment. Cited pages in our set average 844 words against 2,119 for uncited ones.
The external evidence points the same way. Ahrefs ran a causal study with 1,885 pages that added structured data against 4,000 controls: null effect in ChatGPT, null in Google AI Mode, and slightly negative in AI Overviews. And Google's guide on optimizing for generative features, published in May 2026, explicitly names five tactics as unnecessary, among them the llms.txt file and the idea that structured data is a requirement. Reddit, to close the point, scores 23 out of 100 on the checklist and is one of the most cited domains in the study.
Read the full analysis in the study, section 07
What changes if this is true
All eight findings point to the same place, and it is not where the industry says to look. It is not a technical checklist, because above 40 points the checklist stops explaining who wins. It is not a page optimized for every model, because they do not agree on even 4% of the URLs they cite. And it is not your site, because no brand domain cleared 10% of the citations in its own category. AI visibility gets built in the third party sources the model already consults: media, video, and the institutional references of the category. Your site is one of those sources, not the main one.
The first thing you can do with this costs nothing: stop measuring only whether you appear, and start measuring where. We captured the Google live dataset with GEO Observer, our Chrome extension. It is free, has no usage limit, and captures the same surfaces we used here. If you want to see what AI is citing for you, and in what slot, install GEO Observer.
Eight findings, five hypotheses and 26 references, in 20 pages.
Dashcrab GEO Citation Research 2026. Data cutoff: August 11, 2026. The 21,571 citations come from two datasets we never mix: 10,618 from monitoring queries across five models and 10,953 from real Google captures, which share only 307 URLs with each other. The 518 search experiment is reported separately and does not add to the total. Citations come from monitoring programs active on the platform and from field work with GEO Observer. No brands, own domains or individual categories are identified: every figure is reported aggregated or anonymized. The study is observational and does not demonstrate causality in any direction. The full methodology is in the PDF, which is published in Spanish. The terms GEO and AEO come from the Princeton, Georgia Tech, Allen Institute and IIT Delhi paper and from SEO jargon, respectively.