All resources
Blog

Seven things the data actually says about getting cited by AI

Seven evidence-backed findings on third-party mentions, search behavior, schema, llms.txt, content scaling, citation signals, and measurement reliability.

Ask five people how to get cited by ChatGPT and you'll get five confident, mostly overlapping answers: publish more, add schema, write an llms.txt file, keep everything fresh. Some of that turns out to be right. Some of it, when you actually go looking for the study behind the claim, turns out to be closer to folklore. Below are seven findings we traced back to their original source rather than a summary of a summary, each with a large enough sample size to actually mean something, and each with a genuinely practical implication for where to spend your time.

A word on how these were chosen. Each finding below comes from a study with a sample size in the thousands or higher, a stated methodology we could actually check, and, where possible, independent replication by a second research team using a different dataset. Where a finding came from only one source, or from a source with an obvious incentive to shade the result a certain way, we've said so plainly rather than presenting it as settled fact. Where an earlier version of this piece got a number slightly wrong, that's corrected below too.

1Third-party sources drive the large majority of brand mentions

AirOps, a content operations company, analyzed 21,311 brand mentions across ChatGPT, Claude and Perplexity and found that 85% of those mentions came from third-party sources, meaning reviews, forums, comparison articles and independent publications, not the brand's own website.1 A brand was roughly 6.5 times more likely to be mentioned through someone else's content than through its own. That figure isn't a lone outlier either. Muck Rack, a media intelligence company, has run this same measurement three times on a completely separate dataset built from tens of millions of AI-cited links, and earned media has held between 82% and 89% of all AI citations across all three editions, most recently 84% in a May 2026 report covering more than 25 million links.2 Two unrelated organizations, two different methodologies, converging in the low-to-mid 80s is the kind of agreement that's hard to dismiss as coincidence.

85%
of brand mentions in AI search came from third-party sources, per AirOps1
84%
of AI citations traced to earned media, independently, per Muck Rack (May 2026)2
6.5x
more likely to be mentioned via a third party than your own site1

A separate, much larger correlational study adds weight to the same conclusion. Ahrefs analyzed 75,000 brands and found that branded mentions across the web correlate with AI visibility at roughly 0.664, compared with just 0.218 for backlinks, historically the strongest authority signal in classic SEO.3 That's not a small gap. It suggests AI systems are weighing what other people say about you far more heavily than who links to you.

The practical implication cuts against a reflex a lot of marketing teams have: if your brand has a visibility gap in AI answers, the instinct to publish more content on your own domain is treating the wrong symptom. The data suggests the fix lives largely off your site, in the reviews, forum threads and independent coverage that AI systems are actually pulling from. That doesn't mean your own site is worthless. Owned content still needs to be accurate and clear, since it's often the source third parties draw on when they write about you in the first place. It means owned content alone, without any independent validation, is working with only a minority share of the actual opportunity.

2ChatGPT users still lean on Google, though "99%" needs a footnote

Google's own consumer insights team published a headline figure worth taking seriously precisely because it undercuts their own AI narrative rather than flattering it: 99% of people who go to ChatGPT to research a product or service still consult Google Search at some point in that same purchase journey.4 Worth being precise about where that number actually comes from, since it's easy to misstate: it's a Google-commissioned Ipsos survey of 12,038 people identified as "ChatGPT commercial users," fielded in April 2026 across 25 markets, not weighted to reflect population size. Separately, a different wave of the same research program, an Ipsos Global Consumer Journeys survey of AI users versus non-users in December 2025, found shoppers who used AI platforms interacted with about 2.8 times more touchpoints on their path to purchase (12.1 versus 4.2) than shoppers who didn't.5 Put together: AI isn't shortening the path to purchase. It's adding a step, not replacing one.

It's fair to hold the 99% figure with a bit of healthy skepticism about the messenger. It's Google publishing data that happens to reassure Google's own advertisers that Search still matters, and independent journalism examining the wider guide this stat came from has flagged other places where Google's own headline framing outran what the underlying survey footnotes actually supported.6 The good news is you don't have to take Google's word for the core claim. Similarweb, an independent analytics firm with no stake in the answer, found 95% of ChatGPT's user base also used Google in the same month, a figure that held steady even as visits to generative AI platforms grew 70% year over year.7 SparkToro's analysis of desktop clickstream panel data puts the scale gap in context too: Google still handled about 73.7% of US desktop searches in late 2025, versus 2.86% for ChatGPT.8

None of this means AI search has no effect on classic search behavior. Pew Research Center tracked 68,879 real Google searches from a panel of 900 US adults and found that when an AI summary appeared in the results, people clicked a traditional link only 8% of the time, versus 15% when no summary appeared, and ended their search session entirely without any click on 26% of AI-summary searches versus 16% otherwise.9 So classic search isn't disappearing, but it is being used differently: fewer clicks per search, alongside continued heavy reliance on Google as a destination.

AI is adding a step to the purchase journey, not removing one. Treat GEO as additive to SEO, not a replacement for it.

3Adding schema to an already-successful page barely moves AI citations

This one runs against a piece of near-universal GEO advice. Ahrefs ran a matched, difference-in-differences study, the more statistically rigorous of the designs used in this space, tracking 1,885 pages that added JSON-LD schema against a large set of similar pages that didn't, measuring citation counts 30 days before and after the change. The results: Google AI Mode citations moved about 2.4%, ChatGPT about 2.2%, both statistically indistinguishable from random noise, and Google AI Overview citations fell by roughly 4.6%, a decline Ahrefs' own analysis found unlikely to be chance.10

RETHINK: schema is not the citation lever most content teams assume it is

The important caveat, which Ahrefs flagged themselves, is that every page in that dataset was already receiving significant AI Overview citations before the schema went in, and both the treated and control groups were already trending downward before the change, with the treated pages declining just slightly faster. Ahrefs stops short of blaming schema itself for the dip. The study can't tell you whether schema helps a brand-new or thin page become eligible for citation in the first place either. It can only tell you that adding it to a page that's already winning doesn't meaningfully push that page further ahead.

4llms.txt shows no measurable relationship to citations, in four separate studies

Few proposed "AI SEO" tactics have spread faster than adding an llms.txt file, a simple text file meant to point AI crawlers toward a site's most important content. SE Ranking tested whether it actually works by analyzing roughly 300,000 domains, using both statistical correlation tests and a machine-learning model to see whether having the file predicted citation frequency. It found none. Removing the llms.txt variable from their predictive model actually improved its accuracy, meaning the file was adding noise, not signal.11

Three more independent checks landed on the same answer. Trakkr Research ran a separate analysis across 37,894 domains and found, in their words, zero citation advantage.12 Ahrefs took a different angle entirely, checking not citations but whether the file even gets read: crawling 137,000 sites, they found 97% of published llms.txt files received zero traffic from any bot in May 2026, and of the small share that did get requested, most hits came from AI coding tools, not the search-facing assistants that would actually generate a citation.13 OtterlyAI ran a live server-log experiment over 90 days and found only about 0.1% of AI bot visits ever requested the file, versus roughly 265 bot visits per normal content page on the same sites.14

StudyWhat it measuredFinding
SE Ranking~300,000 domains, citation correlationNo measurable lift
Trakkr Research37,894 domains, citation correlationZero citation advantage
Ahrefs137,000 domains, whether the file gets read97% never fetched by any bot
OtterlyAI90-day live server logs~0.1% of AI bot visits requested it

Four independent teams, four different methods, all pointing the same direction: nobody is reading the file, and having it doesn't correlate with getting cited more.

SKIP: llms.txt is low effort, but don't expect it to move citations today

5Scaling AI-generated content produces a boom, then a worse bust

SEO researcher Lily Ray tracked more than 220 websites and subfolders that had been publicly named as customers of AI content scaling platforms, cross-referencing their organic traffic history against Ahrefs and Sistrix data. The pattern she found was strikingly consistent: rapid growth in published pages over six to twelve months, an organic traffic peak roughly three to six months after the content volume peaked, then a steep decline that, in many cases, erased most of the gain and dropped traffic below where it started.15

54%
of the 220+ sites lost 30% or more of their peak organic traffic15
39%
lost 50% or more of their peak traffic15
22%
lost 75% or more, essentially erasing the entire program's gains15

Ray is careful about the limits of her own analysis, and it's worth repeating those limits rather than glossing over them: correlation across 220 sites doesn't prove any single AI content tool caused any single site's decline, since algorithm updates, site migrations, and unrelated business changes can produce similar patterns. What makes the pattern hard to dismiss is its consistency across more than a dozen different vendors and hundreds of sites, with timing that lines up closely with known Google quality-update windows, and a detail that's almost darkly funny: many of the declines began shortly after the vendor's own case study, celebrating the site's traffic growth, was published.

The mechanism behind the pattern isn't mysterious once you see it laid out. A site publishes hundreds of AI-generated pages quickly, often thin variations on the same topic (comparison pages, glossary entries, location pages) built for coverage rather than genuine usefulness. Search engines initially index and rank the new volume, producing a real traffic bump. Over the following months, as Google's quality systems accumulate enough signal about the pattern (thin content, low engagement, high similarity across pages), rankings correct, sometimes sharply and all at once. Because AI answer engines lean on the same underlying web index and retrieval signals, a site that loses organic visibility this way tends to lose AI citations at the same time, not just Google rankings.

CAUTION: rapid, undifferentiated content scaling has a well-documented failure pattern

6What actually does move citations: evidence, quotes and freshness

Everything above is mostly a list of things that don't work. It's worth being equally rigorous about what does, since the evidence here is unusually strong: it comes from a peer-reviewed academic paper rather than a vendor blog. Researchers from Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi formally coined the term "Generative Engine Optimization" in a paper accepted at KDD 2024, one of the top conferences in the field, and tested nine specific content changes against a 10,000-query benchmark to see which ones actually increased a source's chance of being cited.16 Adding statistics, adding direct quotations from credible sources, and citing other authoritative sources each produced meaningful, measurable lifts, in some cases up to about 40%. Keyword stuffing, the classic SEO trick with nothing to do with genuine trustworthiness, produced flat or negative results.

+40%
visibility lift from citations, quotes and statistics, KDD 2024 peer-reviewed study16
25.7%
fresher, on average, is AI-cited content compared to classic organic results, across 17M citations17
~0
lift from keyword stuffing, the same KDD study found

Freshness backs the same story from a different angle. Ahrefs analyzed nearly 17 million citations across ChatGPT, Perplexity, Gemini, Copilot and AI Overviews and compared them against classic organic search results, finding AI-cited pages run about 25.7% fresher on average, with the gap strongest on ChatGPT.17 The honest caveat: the average AI-cited page is still roughly three years old, so freshness is a real factor, not a substitute for depth or authority.

DO: evidence density, real quotes, cited sources, and a genuine freshness signal

7Why no single measurement of "AI visibility" can be trusted

This is the finding that puts all six above in context. SparkToro, working with Gumshoe.ai, had 600 volunteers run 12 identical brand-recommendation prompts through ChatGPT, Claude and Google's AI systems, nearly 3,000 runs in total. The odds of getting the exact same list of recommended brands twice were under 1 in 100. The odds of getting the same list in the same order were closer to 1 in 1,000.18 Separately, Ahrefs found that Google's own AI Mode and AI Overviews cite different sources for the same query 87% of the time, and AirOps found 68% of brands they tracked appeared on only one AI platform, never showing up consistently across two or more.19

The nuance that keeps this from being pure nihilism: the specific brands under consideration tend to stay fairly stable across runs, it's mainly the order and exact selection that shifts. A well-known brand in a category will usually keep appearing somewhere in the results. What won't hold still is any single-prompt "ranking," which is exactly the kind of number a screenshot-friendly dashboard tends to feature.

The common thread across all seven

Put next to each other, these seven findings point in a surprisingly consistent direction. The tactics that looked like quick technical wins, schema markup on already-strong pages, an llms.txt file, publishing content faster than a normal editorial process would allow, all failed to move the needle or actively backfired. The things that did work, third-party validation, genuine evidence density, and real freshness, are slower and harder, and can't be checked off a technical to-do list in an afternoon. And underneath all of it sits a harder truth: even when you do the right things, any single check of "are we visible" is close to worthless without repeating it many times across engines, because the underlying system is genuinely probabilistic.

There's also a broader lesson here about how to read any single GEO claim you come across, including the ones above. A number attached to one small study, one vendor's customer list, or one company with an obvious reason to want a particular answer to be true, deserves more scepticism than a number several separate teams landed on independently. The llms.txt finding is convincing not because SE Ranking is a uniquely trustworthy source, but because three other teams, using different data and different methods, checked the same question and got the same answer. That's the standard worth holding every claim in this space to, including the ones in this article, one of which we corrected after digging into the primary source ourselves.

Track what's actually moving your citations

Ripplix samples every tracked prompt repeatedly across engines, not once, and tracks your brand's third-party mentions, citations and AI-readiness signals in one place.

Get your free AI Visibility Report →
Sources:
  1. AirOps, analysis of 21,311 brand mentions across ChatGPT, Claude and Perplexity, March 2026.
  2. Muck Rack, "What Is AI Reading?" report series, three editions (July 2025, December 2025, May 2026); most recent edition covers 25M+ links across 17 industries.
  3. Ahrefs, analysis of 75,000 brands comparing branded mention correlation (0.664) versus backlink correlation (0.218) with AI visibility.
  4. Google (Think with Google), "4 AI holiday shopping shifts," citing a Google-commissioned Ipsos survey of 12,038 ChatGPT commercial users across 25 markets, April 2026.
  5. Ipsos Global Consumer Journeys survey, cited via Google's consumer insights publications, comparing 14,305 AI users to 21,881 non-users, December 2025.
  6. Independent review of Google's AI shopping guide methodology and footnotes, PPC Land, 2026.
  7. Similarweb, audience overlap analysis between ChatGPT and Google users, reported via Search Engine Journal, 2026.
  8. SparkToro, analysis of Datos desktop clickstream panel data, US search engine share, late 2025.
  9. Pew Research Center, behavioral tracking of 68,879 Google searches from a 900-person US panel, July 2025.
  10. Ahrefs, matched difference-in-differences study of 1,885 pages adding JSON-LD schema, May 2026.
  11. SE Ranking, analysis of llms.txt adoption and AI citation frequency across ~300,000 domains, 2025–2026.
  12. Trakkr Research, independent replication across 37,894 domains, 2026.
  13. Ahrefs, crawl analysis of llms.txt fetch activity across 137,000 domains, June 2026.
  14. OtterlyAI, 90-day live server-log experiment on llms.txt bot requests, 2026.
  15. Lily Ray, independent analysis of 220+ sites publicly named as AI content platform customers, cross-referenced against Ahrefs and Sistrix data, published via Substack and covered by Search Engine Journal, May 2026.
  16. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan & Deshpande, "GEO: Generative Engine Optimization," Princeton University / Georgia Tech / Allen Institute for AI / IIT Delhi, presented at ACM SIGKDD (KDD) 2024.
  17. Ahrefs, analysis of nearly 17 million AI citations across ChatGPT, Perplexity, Gemini, Copilot and AI Overviews, comparing freshness to classic organic results.
  18. SparkToro and Gumshoe.ai, repeated-prompt study across 600 volunteers and nearly 3,000 runs on ChatGPT, Claude and Google AI, January 2026.
  19. Ahrefs, analysis of source overlap between Google AI Mode and AI Overviews; AirOps, cross-platform brand mention consistency analysis.