Dashcrab

SEO & AI visibility

Which GEO tactics work, and which ones are myth?

Four findings replicated across different samples, six beliefs that do not survive a controlled experiment, and two open hypotheses.

Four things about GEO have replicated across different samples, by teams that do not talk to each other. And six beliefs the industry repeats do not survive a controlled experiment. This article separates the two lists, with the sources.

If you are starting from zero, begin with what GEO and AEO are. Here we assume that vocabulary: citation, mention, retrieval.

What is reasonably established

These are not certainties. They are the closest thing we have.

1. Retrieval rules over writing

The foundational GEO paper tested nine content modifications over about 10,000 queries. Adding verifiable statistics, quotations and sources improved visibility by up to 40%. Keyword stuffing did nothing. That is the headline everyone repeats.

What almost nobody tells you is the counter-experiment. Puerto, Gubri, Green, Oh and Yun published C-SEO Bench in 2025 and tested nine methods, seven taken from the original paper, over six domains and 1,921 queries. Most turned out to be barely effective.

The point of comparison is devastating: moving the document to the first position of the retrieved context was, in retail, about 7.6 times more effective than the best writing tactic. Translated: entering retrieval well matters more than writing pretty for generation.

Our own data points the same way. We crossed 1,486 URLs with a known AEO score against their real citations: not a single page below 40 points out of 100 received a single citation. And above that threshold the score stops ordering anything: C-grade pages were cited seven times more than B pages. There is a step, not a staircase.

Open hypothesis 01

The technical checklist works as an entry requirement, not as a competitive advantage. If that is true, going from 20 to 50 points buys you the right to play and is cheap. Going from 70 to 95 is the most expensive spend and the one with the least evidence.

2. Each engine eats from a different plate

Overlap between models in our citation set comes out at 3.5% at the exact-URL level and 10.4% at the domain level. Three or four pages in common out of every hundred.

It is not a quirk of ours. Writesonic, over 161,286 queries, found 17% shared sources. Profound reported 89% distinct citations between ChatGPT and Perplexity. And the work of Chen, Wang, Chen and Koudas, from the University of Toronto, documents that engines differ in domain diversity, freshness and stability across languages.

In practice: Claude sends 42.5% of its citations to company blogs and ChatGPT barely 12.6%. Google AI Mode cites social and video in 12.1% of cases and Claude in 0.5%. That is not a percentage adjustment; it is a different content playbook.

3. Citations barely generate clicks

Pew Research Center tracked the real behavior of 900 adults in the United States, over 68,879 searches. With an AI summary on screen, clicks on a traditional result fell from 15% to 8%. A click on a source cited inside the summary happened in 1% of visits.

If almost nobody clicks, what remains is the impression: that your name is there, and in what position. In our set, the blog is the most cited page type by volume, with 8,139 citations, and it reaches first position only 7.5% of the time. Reference sites have twenty times less volume and reach first place 24.7% of the time.

4. Your site is a minority in your own category

When we count what share of a category's citations points to the site of a brand in that same category, no brand domain exceeded 10%. The average was 3.3%. The rest goes to media, forums, institutions, platforms and competitors.

Watch out for a confusion that causes panic in teams doing things right. Several studies report that between 38% and 70% of citations go to brand sites. That number and ours are both correct and measure different things: brand sites is the sum of every company in the world, your competitors included.

Full study

This is where the first-party data comes from: 20 pages, every table and all 26 references.

Read the full study

What gets repeated and does not survive the data

What people believe What the evidence shows
Schema markup gets you cited more Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 controls: +2.4% in AI Mode, +2.2% in ChatGPT, both indistinguishable from zero, and -4.6% in AI Overviews, which is significant and points the other way
llms.txt is the entry door Google, May 2026: it can be crawled like any text file and gets no special treatment. SE Ranking analyzed close to 300,000 domains and found no association with citation frequency
You have to fragment content for AI Google says its systems understand pages covering several topics and extract the relevant passage without pre-chunking
You can optimize one page for every LLM Exact-URL overlap is 3.5% in our data, and no independent study has put it above 20%
Longer and more exhaustive gets cited more Cited pages in our set average 844 words. Uncited ones, 2,119
Ads help you appear in the answer They were 0.03% of AI Overview citations in our 518 captured searches, and none of the 75 paid pages survived the move to AI Mode

The schema case deserves a paragraph because it is the best free lesson in the space. Ahrefs had earlier found, over 6 million URLs, that AI-cited pages carry JSON-LD almost three times as often. That correlation circulated for two years as proof.

When they ran the controlled experiment, the effect evaporated. The explanation is boring: sites that implement structured data also do technical SEO, publish better and earn links. Schema was riding along as a stowaway. A parallel experiment by searchVIU also found that five AI systems, when fetching a page live, extract only the visible HTML and do not read the hidden markup.

None of this means removing schema. It works for rich results, knowledge graphs and entity disambiguation. It means it stopped being a citation lever.

Open hypothesis 02

A good part of the GEO checklist is not a cause of citation; it is a symptom of well-maintained sites. Every new tactic that appears should be run through the same question: does this move anything, or is it simply what sites that were going to be cited anyway already do?

Before moving budget on the next industry headline, there are five questions worth asking any study, including ours. They are in how to read a GEO study without swallowing the headline. And if you want to see what AI is citing about you, GEO Observer is free and captures your real searches.


First-party figures come from Dashcrab GEO Citation Research 2026, with a data cutoff of August 11, 2026, and from the experiment of 518 searches captured with GEO Observer between July and August 2026. They are reported aggregated or anonymized, and our studies are observational: they do not demonstrate causality. Third-party figures are linked to their original source. The full methodology is in the study PDF.

Frequently asked questions

Is structured data worth implementing?

Yes, but not for what it gets sold for. The controlled Ahrefs experiment on 1,885 pages found no increase in citations on any platform, and Google says there is no special markup for generative features. Keep using it for rich results and entity disambiguation, and take it off your list of citation levers.

Should I create an llms.txt file?

For Google it changes nothing: it gets crawled like any text file, with no special treatment. It can make sense if a specific service you care about consumes it. As a general visibility tactic there is no evidence behind it.

Does longer content get cited more?

Not in our data. Cited pages in our set average 844 words and uncited ones 2,119. The likely explanation is that a model extracts a short, specific passage; it does not reward exhaustiveness. Write passages that answer a concrete question in the first two sentences.

Is there a tactic that works across every AI model at once?

Almost none at the page level: exact-URL overlap between models is 3.5% in our data, and no independent study has put it above 20%. What they do share is the floor: retrieval, clean URLs and self-sufficient passages. Beyond that, each model rewards different formats and you have to measure model by model.