Dashcrab

SEO & AI visibility

Can you trust a GEO study?

What the field still does not know, and the five questions to ask any study before moving budget on a headline. They apply to ours too.

Most GEO headlines do not survive five questions. This article gathers what the field still does not know, which is the part almost no content in this space writes, and the five questions we ask any study before moving a dollar. Ours included.

What remains open

How much of this is causal. Very little. The available corpus, ours included, is overwhelmingly observational: you look at citations that already happened and search for patterns. The two serious causal experiments that exist, C-SEO Bench and the Ahrefs schema study, produced null results or results contrary to expectations. That does not prove nothing works. It proves the industry is running ahead of its own evidence.

How much of what you see in a panel is signal and how much is noise. SISTRIX analyzed three platforms, six countries and seventeen weeks: at the domain level, AI Mode rotates 56% weekly and ChatGPT 74%. At the URL level rotation reaches 85%.

SE Ranking ran 10,000 keywords three times on the same day in AI Mode: the average overlap of exact URLs between the three runs was 9.2%. If your monthly report says you went up, the first honest question is whether that sits above the system's own noise.

How much an AI click is worth. Vendor reports agree on direction and not on magnitude: Semrush reported 4.4 times more conversion than organic, Seer Interactive measured 15.9% against 1.76%, Ahrefs counted 0.5% of its sessions generating 12.1% of its signups. When five measurements of the same phenomenon range from 4 to 23 times, you are not looking at a benchmark; you are looking at an uncertainty range.

What happens when everyone optimizes. C-SEO Bench showed diminishing returns with adoption. And Kumar and Lakkaraju, from Harvard, demonstrated something more uncomfortable: with an adversarially designed text sequence they pushed a product to the top of a model's recommendations. If that scales, the platforms' response will be the same one Google had to link spam.

What happens outside English. Almost all the public literature ran in English and in the United States. How much applies to a query in Rioplatense Spanish about a local category is, today, an unanswered question. It is the one that interests us most and the one we are measuring.

The five questions

Which surface did it measure? AI Overview, AI Mode, ChatGPT and Perplexity are different systems. In our 518 captured searches, 47.45% of AI Overview citations were already in the organic results of that same search, and on moving to AI Mode only 19.98% survive. That is why headlines saying 76% of citations come from the top 10 coexist with others saying 17%. They do not contradict each other: they measure different screens.

Is it observational or causal? If nobody changed anything on purpose and compared against a control group, it is a correlation. It works for generating hypotheses, not for justifying budget.

Who asked the questions? Synthetic queries written by a team are not real user queries. Both are useful, for different things.

Does it measure appearance or position? Appearing and appearing first are metrics that move in opposite directions. The company blog is the best example: it wins on volume and loses on position.

In what language and in what market? If it does not say, assume English and the United States.

The first question of your own study you can answer today, free. GEO Observer captures your real searches in ChatGPT, Gemini and Google, with no usage limit. And continuous tracking, with a weekly scan of more than 100 questions, lives in AI Visibility.


First-party figures come from Dashcrab GEO Citation Research 2026 and from the experiment of 518 searches captured with GEO Observer between July and August 2026. Our studies are observational and do not demonstrate causality. Third-party figures are linked to their original source.

Frequently asked questions

How long until a GEO change shows up?

There is no good answer yet, and that is the honest answer. With weekly rotations of 56% in AI Mode and 74% in ChatGPT at the domain level, any reading shorter than several weeks is measuring noise as much as effect.

How much is a click from an AI answer worth?

Nobody knows precisely. Vendor reports range from 4 to 23 times the conversion of organic traffic, a range too wide to be a benchmark. On top of that, a good part of AI-originated traffic arrives as direct or as branded search, so attribution is undercounted.

What is the difference between an observational study and a causal one?

An observational study looks at citations that already happened and searches for patterns: it works for generating hypotheses. A causal one changes something on purpose and compares against a control group: it is the only kind that can justify budget. Almost everything published about GEO, our studies included, is observational.

Can I trust the report from my AI visibility tool?

Yes, but read the baseline variation first. At the domain level, AI Mode rotates 56% weekly and ChatGPT 74%, and the same run repeated on the same day matches on 9.2% of URLs. Any rise or fall has to beat that noise before it counts as signal.