Most GEO headlines do not survive five questions. This article gathers what the field still does not know, which is the part almost no content in this space writes, and the five questions we ask any study before moving a dollar. Ours included.
What remains open
How much of this is causal. Very little. The available corpus, ours included, is overwhelmingly observational: you look at citations that already happened and search for patterns. The two serious causal experiments that exist, C-SEO Bench and the Ahrefs schema study, produced null results or results contrary to expectations. That does not prove nothing works. It proves the industry is running ahead of its own evidence.
How much of what you see in a panel is signal and how much is noise. SISTRIX analyzed three platforms, six countries and seventeen weeks: at the domain level, AI Mode rotates 56% weekly and ChatGPT 74%. At the URL level rotation reaches 85%.
SE Ranking ran 10,000 keywords three times on the same day in AI Mode: the average overlap of exact URLs between the three runs was 9.2%. If your monthly report says you went up, the first honest question is whether that sits above the system's own noise.
How much an AI click is worth. Vendor reports agree on direction and not on magnitude: Semrush reported 4.4 times more conversion than organic, Seer Interactive measured 15.9% against 1.76%, Ahrefs counted 0.5% of its sessions generating 12.1% of its signups. When five measurements of the same phenomenon range from 4 to 23 times, you are not looking at a benchmark; you are looking at an uncertainty range.
What happens when everyone optimizes. C-SEO Bench showed diminishing returns with adoption. And Kumar and Lakkaraju, from Harvard, demonstrated something more uncomfortable: with an adversarially designed text sequence they pushed a product to the top of a model's recommendations. If that scales, the platforms' response will be the same one Google had to link spam.
What happens outside English. Almost all the public literature ran in English and in the United States. How much applies to a query in Rioplatense Spanish about a local category is, today, an unanswered question. It is the one that interests us most and the one we are measuring.
The five questions
Which surface did it measure? AI Overview, AI Mode, ChatGPT and Perplexity are different systems. In our 518 captured searches, 47.45% of AI Overview citations were already in the organic results of that same search, and on moving to AI Mode only 19.98% survive. That is why headlines saying 76% of citations come from the top 10 coexist with others saying 17%. They do not contradict each other: they measure different screens.
Is it observational or causal? If nobody changed anything on purpose and compared against a control group, it is a correlation. It works for generating hypotheses, not for justifying budget.
Who asked the questions? Synthetic queries written by a team are not real user queries. Both are useful, for different things.
Does it measure appearance or position? Appearing and appearing first are metrics that move in opposite directions. The company blog is the best example: it wins on volume and loses on position.
In what language and in what market? If it does not say, assume English and the United States.
The first question of your own study you can answer today, free. GEO Observer captures your real searches in ChatGPT, Gemini and Google, with no usage limit. And continuous tracking, with a weekly scan of more than 100 questions, lives in AI Visibility.
First-party figures come from Dashcrab GEO Citation Research 2026 and from the experiment of 518 searches captured with GEO Observer between July and August 2026. Our studies are observational and do not demonstrate causality. Third-party figures are linked to their original source.