Dashcrab

SEO and AI visibility

What pages does AI actually cite when it answers?

We analyzed 21,571 URL citations from Google AI Overview, ChatGPT, Perplexity, Claude, Google AI Mode and Gemini. Eight findings, and five places where our data does not match what the industry keeps repeating.

We analyzed 21,571 URL citations from six AI models and found eight things. A citation is every link a model shows as a source for its answer: if one answer lists seven sources, we count seven. The most uncomfortable of the eight: no brand site cleared 10% of the citations in its own category. The most useful one: the typical citation is not your homepage, it is a specific internal page. And the one that separates us most from what is already published: blog content is the most cited type by volume and second to last by position inside the source list. This article walks through the eight findings with their strongest data point, and each one links out to the full analysis.

21,571
URL citations analyzed
14,414
Unique URLs
7,282
Distinct domains competing
6
AI models analyzed
Full study, in Spanish

All 20 pages, with every table and the 26 references.

Read the full study

The typical citation is not your homepage

01. Models cite internal pages that are clean and specific

Between 67.8% and 75.2% of citations point to pages sitting two or more levels deep inside the site. Not example.com and not example.com/blog, but example.com/blog/how-to-measure-geo. 99.98% of citations use HTTPS, 94.1% do not end in a file extension, and only 4.2% carry parameters in the address.

This separates two things people tend to conflate: a model citing your company, or a model citing a specific answer you published. The second is far more common, and that changes where the investment belongs.

What makes the finding strong is not the percentage, it is the repetition. The spike at depth 2 shows up just as sharply in both datasets, and those two datasets share only 3.7% of their URLs. It is not an artifact of one instrument or one category. The implication is boring and hard to sell: a good chunk of AI readiness work is architecture hygiene. Clean, stable, specific addresses. There is no trick here, there is an entry requirement plenty of sites still do not meet.

Read the full analysis in the study, section 02

02. Blogs win volume and lose position

When a model answers, it lists its sources in an order, and that order is not cosmetic: the first one is the most visible and, in several interfaces, the only one you see without expanding the rest. Almost no published study measures that. They measure whether you showed up. We added where, and the picture changes.

Page type Citations Reaches 1st slot Average position
Homepage1,16514.2%5.9
Guide or resource2,43612.5%8.4
Social and video1,83511.4%7.5
Product page4,8418.6%8.5
Company blog8,1397.5%10.2
PDF4116.8%7.4

The company blog has twenty times the volume of reference sites and less than a third of their first slot rate. It is the most cited content type in the study and among the last in quality of placement. And placement matters more than it looks, because clicks are vanishingly rare: Pew Research found that only 1% of users click any link inside an AI Overview. If nobody clicks, the value of a citation concentrates in how visible its slot is.

Dashcrab hypothesis 01

Blogs dominate AI citations and blogs win AI citations are two different claims, and the industry uses them as synonyms. We propose splitting reach (do we show up?) from quality of placement (do we show up first?). Blogs optimize the former and are mediocre at the latter. This does not say stop blogging. It says: stop measuring mentions alone, and do not expect the blog to hand you the top slot, because structurally it does not.

Read the full analysis in the study, section 03

Four models mostly cite the blog. ChatGPT does not

03. There is no such thing as one URL for every LLM

Any combination of models we take agrees on barely 3.5% of the exact URLs it cites. At the domain level the overlap rises to 10.4%. Put another way: they may trust the same sites, but they rarely pick the same page.

Model Citations Blog Product Homepage Social and video
Claude2,90942.5%28.2%6.9%0.5%
Google AI Mode2,51439.1%23.4%5.6%12.1%
ChatGPT2,15112.6%31.3%10.6%0.7%
Gemini1,56738.7%22.4%16.5%3.1%
Perplexity1,47733.1%30.6%7.0%4.7%

Classification by page type and site type over the 10,618 citations of the multi model set. 97% of the citations were identified.

Four of the five mostly cite the company blog, at weights ranging from 33.1% to 42.5%. ChatGPT does not: its most cited category is product pages, at 31.3%, while its blog reaches only 12.6%. Add guides, PDFs and documentation and 62% of what ChatGPT cites is product material.

The other wide gap is social, and we now know what it is made of: Google AI Mode runs at 12.1%, split into 6.8% social networks, 3.6% video and 1.7% LinkedIn. Claude adds up to 0.5% across all three. A twenty-fourfold difference. That is not a percentage tweak, it is a different content priority.

The commercial consequence is finer than it sounds: you cannot promise the same URL wins on every LLM. There is shared ground, especially on product pages. What you build is a portfolio: types that serve several models, formats where you specialize, and measure model by model.

What each model cites, one by one

04. Nobody owns the conversation, and your site least of all

The single most cited domain in the entire study accounts for 3.6% of volume. Not even YouTube, which leads it, reaches 4%. You need 500 domains to cover half the citations. There is no oligopoly: the long tail is the market, with 7,282 domains competing.

Now the number that matters. When we measure how many citations in a category point at the site of a brand in that same category, no brand domain cleared 10%, and the average was 3.3%. The remaining 96.7% goes to third parties: media, forums, institutions, platforms and competitors. Your brand is not in the conversation. It is being mentioned by others, or it is not there.

Dashcrab hypothesis 02

The industry reports that brand sites account for between 38% and 70% of citations, depending on the model. That number and ours are both correct, and they measure different things. Brand sites is simply the sum of every commercial site in the world: a company, its competitors, and any business relevant to the question. That half of ChatGPT citations go to corporate sites says nothing about how many go to yours. If someone hears 50% and their dashboard shows 4%, the natural conclusion, and the wrong one, is that their team is failing.

The work of Chen, Wang, Chen and Koudas, at the University of Toronto, documents the same thing from another angle: a systematic bias in AI search engines toward authoritative third party sources, over brand owned content. The practical consequence is not to give up. It is to stop treating your own site as the primary AI visibility channel and start treating it as one among several, alongside presence in media, in video, and in the institutional sources of the category.

Read the full analysis in the study, section 05

All eight findings, in Spanish

With the methodology, the tables by industry, and what the study does not say.

Read the full study

The two faces of Google

Google has two surfaces that behave very differently and that were being measured as if they were one: the AI Overview, the summary that shows up above the results as soon as you search, and AI Mode, the follow-up chat you land in when you ask again. The three findings that follow come from a parallel experiment: 518 Google searches, between July and August 2026, covering ten industries and two markets. For each one we captured all four surfaces from the same screen and the same session, using GEO Observer: the AI Overview, the organic results, the paid ads, and AI Mode on the follow-up. That design is what makes it possible to compute overlap per page within the same search, instead of comparing captures taken at different moments.

05. AI Overview repeats your SEO. AI Mode throws it out

47.45% of AI Overview citations were already in the organic results of that same search. Ask a follow-up and move to AI Mode, and only 19.98% of those organic pages survive. And AI Mode does not just discard: it also multiplies. The average goes from 7.53 citations per AI Overview to 14.98 per AI Mode turn, and 85.29% of those citations were not in the organic results, the ads, or the AI Overview of that search.

This resolves a contradiction the industry has been carrying for two years. When a study reports that 76% or 97% of AI citations come from the organic top 10, it is looking at AI Overviews. When another reports 12% or 17%, it is looking at AI Mode or ChatGPT. They do not contradict each other: they describe different surfaces of the same system, and the label AI citations throws them all in the same bucket.

Read the full analysis of the 518 searches

06. Paid ads do not exist for the generative answer

Ads accounted for 0.03% of AI Overview citations. And persistence into AI Mode was 0% across 75 paid pages, without a single exception in ten industries. The most plausible explanation is one of genre, not platform policy: an ad is written to convert someone who already decided to click, and a citation has to read like a source that answers the question.

The budget implication is concrete. If someone justifies ad spend with the argument of presence in generative answers, this data contradicts them directly.

Read the full analysis of the 518 searches

07. AI Overview overweights video far beyond SEO

YouTube weighs 1.72% in organic results and 7.72% in AI Overview citations. More than four times. Add social media and it reaches 16.08%. Reddit and Wikipedia together, in the same sample, add up to 1.57%. There is a received idea that generative answers favor forums and encyclopedias, and in this data it is the reverse by an order of magnitude.

AI Overview does not inherit the proportions of SEO. It actively overweights video and social content. If your category lends itself to that, a well structured video can weigh more than an article optimized the old way.

Read the full analysis of the 518 searches

AEO has a floor, not a staircase

08. There is a threshold to cross, and then a noisy plateau

An AEO score is a 0 to 100 grade for how ready a page is to be cited. It combines technical, formatting and identity signals. Almost every tool in the space, ours included, computes something similar, and the implicit promise is that raising the score raises the odds of being cited. We crossed 1,486 URLs with known scores against the citations they actually received, and that promise holds only in one stretch.

Page grade Citation rate
A · 85 and up11.9%
B · 70 to 842.4%
C · 55 to 6917.5%
D · 40 to 542.8%
F · under 400%

The floor is absolute: no page below 40 points received a single citation. Above it, the score stops sorting anything. Pages graded C get cited seven times more often than pages graded B, which score better. There is no staircase that climbs with the score. There is one step and then a noisy plateau.

The most uncomfortable data point sits in the formatting component. Ranking articles by their structure and presentation score, the relationship with real citation runs backwards: the worst scoring quartile gets cited 9.1% of the time, and the best scoring quartile, 0%. The likely explanation is not that formatting hurts, but that the score captures something else. Pages that score 100 on structure are long, exhaustive articles, the kind a brand produces to rank, and what a model extracts is a short fragment. Cited pages in our set average 844 words against 2,119 for uncited ones.

The external evidence points the same way. Ahrefs ran a causal study with 1,885 pages that added structured data against 4,000 controls: null effect in ChatGPT, null in Google AI Mode, and slightly negative in AI Overviews. And Google's guide on optimizing for generative features, published in May 2026, explicitly names five tactics as unnecessary, among them the llms.txt file and the idea that structured data is a requirement. Reddit, to close the point, scores 23 out of 100 on the checklist and is one of the most cited domains in the study.

Read the full analysis in the study, section 07

What changes if this is true

All eight findings point to the same place, and it is not where the industry says to look. It is not a technical checklist, because above 40 points the checklist stops explaining who wins. It is not a page optimized for every model, because they do not agree on even 4% of the URLs they cite. And it is not your site, because no brand domain cleared 10% of the citations in its own category. AI visibility gets built in the third party sources the model already consults: media, video, and the institutional references of the category. Your site is one of those sources, not the main one.

The first thing you can do with this costs nothing: stop measuring only whether you appear, and start measuring where. We captured the Google live dataset with GEO Observer, our Chrome extension. It is free, has no usage limit, and captures the same surfaces we used here. If you want to see what AI is citing for you, and in what slot, install GEO Observer.

Download the PDF, in Spanish

Eight findings, five hypotheses and 26 references, in 20 pages.

Read the full study

Dashcrab GEO Citation Research 2026. Data cutoff: August 11, 2026. The 21,571 citations come from two datasets we never mix: 10,618 from monitoring queries across five models and 10,953 from real Google captures, which share only 307 URLs with each other. The 518 search experiment is reported separately and does not add to the total. Citations come from monitoring programs active on the platform and from field work with GEO Observer. No brands, own domains or individual categories are identified: every figure is reported aggregated or anonymized. The study is observational and does not demonstrate causality in any direction. The full methodology is in the PDF, which is published in Spanish. The terms GEO and AEO come from the Princeton, Georgia Tech, Allen Institute and IIT Delhi paper and from SEO jargon, respectively.

Frequently asked questions

What exactly counts as an AI citation?

Every time a model shows a link as a source for its answer. If one answer lists seven sources, that is seven citations. If the same page shows up in two different answers, that is two citations. On top of counting them, we recorded where in the list each one appeared, which is the data point behind the finding about blogs.

Is blogging worth it if it never reaches the first slot?

Yes, but not for the reason people expect. The company blog is the most cited type by volume, with 8,139 citations in our dataset, and it is what gets you into the conversation. What it does not do is give you the top slot: it lands there 7.5% of the time, with an average position of 10.2. Reference sites land first 24.7% of the time with twenty times less volume.

Can I optimize one page for every AI model at once?

The same exact URL, almost never: overlap comes out at 0.035, which is 3 or 4 pages in common out of every 100. At the domain level it rises to 0.104. That does not mean there is no shared ground: product pages weigh around 30% for Claude, ChatGPT and Perplexity. What works is a portfolio: page types that serve several models, formats where you specialize, and measuring model by model.

Do Google Ads help me show up in AI answers?

No. Ads accounted for 0.03% of AI Overview citations, and of the 75 paid pages we captured, none survived the jump to AI Mode. Zero, without a single exception across ten industries. Search ad spend does what it does, but it is not a lever for AI visibility.

How high do I need to push my AEO score?

Just past the floor. No page scoring under 40 out of 100 received a single citation across our set of 1,486 URLs. Above that threshold the score stops sorting anything: pages graded C get cited seven times more often than pages graded B. Taking a site from 20 to 50 is cheap and unlocks the game. Taking it from 70 to 95 is the most expensive effort in the plan and the one with the least evidence behind it.