Ask ChatGPT the same question twice and you will often get two different lists of brands. Ask Google's AI Overviews the same question next week and many of the links underneath the answer will have changed. This is normal behavior for AI search, not a glitch in your tracking.
That creates a real problem for anyone measuring AI visibility. If the answer moves on its own, how do you tell a real loss from random movement? Below is what the research says about how much AI answers change, why, and how to measure it yourself in a spreadsheet.
What citation churn means
- Citation. A link that an AI answer shows as its source. In AI Overviews, ChatGPT search or Perplexity, these are the linked sources shown with the answer.
- Brand mention. Your brand's name appearing in the answer text, with or without a link.
- Citation churn. How much the set of cited links changes when you ask the same question again, either a few minutes later or weeks later.
- Citation persistence. The opposite of churn. It is the share of links that are still being cited after a set amount of time.
So if 100 links were cited for your questions today and 60 of them are still cited in a week, your 7-day persistence is 60% and your 7-day churn is 40%.
Two kinds of change hide inside that number. Run-to-run noise is the AI giving a slightly different answer each time, even when nothing has changed. Real drift is the AI actually moving toward or away from certain sources over time. Good measurement separates the two. Single-run tracking cannot.
What the research says about how unstable AI answers are
Several independent teams have measured this with different methods. They found the same basic pattern: the exact answer moves a lot, while the general meaning moves much less.
Brand lists rarely repeat
SparkToro, an audience research company, ran a study in late 2025 and published it in January 2026. In it, 600 volunteers ran 12 different prompts through ChatGPT, Claude and Google's AI tools a combined 2,961 times.1 The prompts asked for brand or product recommendations in categories like headphones and cloud computing.
The headline result was blunt. Rand Fishkin wrote that "there's a <1 in 100 chance that ChatGPT or Google's AI, if asked 100X, will give you the same list of brands in any two responses."1 The order of the list was even less stable. He estimated it was "more like 1 in 1,000 runs before you'd see two lists in the same order."1 Even the length of the list changed, from two or three recommendations up to ten or more.
But the study also found something steady underneath the noise. Some brands showed up again and again. In ChatGPT's answers to a question about the best cancer care hospital on the US West Coast, City of Hope "showed up in 69/71 answers: a 97% visibility rate."1 In the headphones test, Bose, Sony, Sennheiser and Apple appeared in 55% to 77% of answers, even though volunteers wrote their prompts in very different ways.1
This means two things. Rank position in an AI answer tells you very little. But how often a brand appears, measured across many runs, can be useful.
Cited links change over days and months
The links underneath AI answers are also unstable. Here are four independent measurements.
| Study | What was measured | Key result |
|---|---|---|
| Max Planck Institute for Software Systems and partners (arXiv paper, 4,706 queries) | Google AI Overviews links, same queries run about two months apart (July/August and September 2025) | Only 18% of web pages were common to both runs. Organic Google results kept 45%.2 |
| Semrush (3,000 keywords, October 2024) | URLs in AI Overviews checked daily for a month | No keyword kept 100% the same URLs. Only 6.7% of desktop URLs seen in the first week were still there in the final week (4.7% on mobile).3 |
| Ahrefs (43,000+ keywords, one month) | Consecutive AI Overview snapshots for the same keyword | On average only 54.5% of URLs overlapped between one snapshot and the next. That means 45.5% of cited sources were new each time.4 |
| Ahrefs (540,000 query pairs, September 2025, US) | AI Overviews compared with AI Mode for the same query | Only 13.7% of citations overlapped between the two.5 |
Semrush found that 91% of URLs were removed from an AI Overview at some point during the month. Only 43% of those removed URLs came back later in the month.3 On average, the same URL stayed in place for 3.87 days in a row on desktop and 3.33 days on mobile.3 This study is from late 2024, so the numbers may have changed since.
Ahrefs found that AI Overviews had "a 70% chance of changing from one observation to the next," and that their content changed every 2.15 days on average.4 Named brands and other entities were also unstable: 54% stayed the same when an AI Overview changed.4
The meaning stays steadier than the sources
Here is the important twist. In the same Ahrefs study, the meaning of the answers hardly moved. AI Overviews scored "an average cosine similarity score of 0.95, where 1.0 represents a perfect match."4 Cosine similarity is a way of scoring how close two pieces of text are in meaning. A score of 0.95 means the answers said almost the same thing, even though nearly half of the links behind them had changed.
The AI Mode comparison shows the same pattern. AI Mode and AI Overviews shared only 13.7% of their citations, but their answers had 86% semantic similarity on average.5 In other words, they usually agreed on what to say while pointing to different pages to support it.
For a brand, this matters a lot. Your page can drop out of the citations without the answer itself changing. And a competitor's page can take your link slot while the answer still describes the market in the same way.
Why AI answers change from run to run
- The question gets split into many searches. Google says AI Overviews and AI Mode "may use a 'query fan-out' technique," which means "issuing multiple related searches across subtopics and data sources" to build a response.6 OpenAI's help page says ChatGPT search usually rewrites your question into one or more targeted searches and sends them to its search partners.7 If the rewritten searches differ a little, the pages they return differ too.
- Different Google features use different systems. Google's own documentation says "AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary."6 That helps explain why the two share so few citations.
- The model itself is not fully repeatable. Even with randomness turned down to zero (a setting called "temperature 0"), AI models can give different outputs. Researchers at Thinking Machines Lab asked one open model the same question 1,000 times at temperature 0. They got 80 different completions.8 They traced the main cause to server load. How many other requests a server is handling at once can slightly change the math, and that small change can tip the model into a different word.8
- Small differences grow. In the Thinking Machines test, all 1,000 answers were identical for the first 102 tokens (a token is roughly a word or part of a word). Then they split: 992 said "Queens, New York" and 8 said "New York City."8 In a search answer, one different word early on can lead to a different follow-up search and a different set of links.
- The web changes. Pages get published and updated, so over weeks real drift adds to the run-to-run noise.
Academic work agrees. A study of five language models on eight tasks found accuracy could vary by up to 15% across runs with settings that were supposed to be fixed.9 The authors concluded that none of the models "consistently delivers repeatable accuracy across all tasks, much less identical output strings."9 The Max Planck study found that even at temperature 0, between 9% and 27% of questions with a yes, no or mixed answer flipped to a different answer when asked again five minutes later.2
So you cannot remove this noise. You can only measure it and plan around it.
A worked example: measuring citation persistence and half-life
Illustrative composite, not real data. "Quillmoor Analytics" is an invented software brand. Every number in this section was made up to show the arithmetic. None of it comes from a real client, a real tracking run or a third-party study.
Step 1: Set up the prompts and the schedule
Quillmoor picks 20 prompts that its buyers might ask, such as "best tools for tracking marketing spend for small teams." Some prompts name Quillmoor and most do not.
The team then follows this plan for one AI platform:
- Run each of the 20 prompts 5 times on day 0. That gives 100 answers.
- Repeat the full set of 5 runs per prompt on day 7, day 14 and day 28.
- Record every cited URL, which prompt it appeared on, and whether Quillmoor is mentioned.
Step 2: Build the day-0 baseline
On day 0, the team lists every URL cited at least twice across the 5 runs of the same prompt. Requiring two appearances filters out links that showed up only once by chance.
That leaves 120 baseline URLs across the 20 prompts. Of those, 30 are Quillmoor's own pages and 90 are third-party pages, such as review sites, comparison articles and forums.
Step 3: Measure the noise floor first
Before looking at weeks, the team checks how much the answers move within a single day. A few hours after the first batch, they run the same 20 prompts 5 more times.
108 of the 120 baseline URLs are cited again at least once. So:
108 ÷ 120 = 0.90, which is 90% same-day persistence
This is the noise floor. Even with nothing changing in the world, about 10% of links drop out between two batches on the same day. Any weekly drop has to be clearly bigger than this before it means anything.
Step 4: Count what survives at 7, 14 and 28 days
On each later date, a baseline URL counts as "still cited" if it appears at least once in that day's 5 runs of the same prompt. Persistence is the number still cited divided by 120.
| Day | Baseline URLs still cited (of 120) | Persistence | Churn (100% minus persistence) |
|---|---|---|---|
| 0 (same-day re-run) | 108 | 108 ÷ 120 = 90.0% | 10.0% |
| 7 | 86 | 86 ÷ 120 = 71.7% | 28.3% |
| 14 | 63 | 63 ÷ 120 = 52.5% | 47.5% |
| 28 | 33 | 33 ÷ 120 = 27.5% | 72.5% |
Illustrative composite data for an invented brand. Not real measurements.
Step 5: Estimate a simple half-life
"Half-life" here means the number of days until half of the baseline links are no longer cited. There are two easy ways to estimate it.
Method A: find where the line crosses 50%. Persistence was 52.5% on day 14 and 27.5% on day 28. So it crossed 50% somewhere between those two dates. Assume it fell at a steady pace between them.
- Total drop between day 14 and day 28: 52.5 − 27.5 = 25 points over 14 days
- Drop needed to reach 50%: 52.5 − 50 = 2.5 points
- Share of the 14 days needed: 2.5 ÷ 25 = 0.1
- Days needed: 0.1 × 14 = 1.4 days
- Half-life: 14 + 1.4 = about 15.4 days
Method B: assume a steady rate of decay. This treats link loss like a fixed percentage lost per day, the same way radioactive half-life works. Use the day-28 number.
- Natural log of persistence: ln(0.275) = −1.291
- Daily decay rate: 1.291 ÷ 28 = 0.0461
- Half-life: ln(2) ÷ 0.0461 = 0.693 ÷ 0.0461 = about 15.0 days
The two methods agree closely. Method B with the day-7 and day-14 numbers gives about 14.6 and 15.1 days. Because these are close, the steady-decay idea fits well. If they were far apart, you would report Method A only.
Step 6: Split your own pages from third-party pages
The overall number hides a useful detail. Here is the same data split by who owns the page.
| Group | Baseline | Day 7 | Day 14 | Day 28 | Half-life (Method A) |
|---|---|---|---|---|---|
| Quillmoor's own pages | 30 | 26 (86.7%) | 22 (73.3%) | 16 (53.3%) | Over 28 days (never fell below 50%) |
| Third-party pages | 90 | 60 (66.7%) | 41 (45.6%) | 17 (18.9%) | About 12.5 days |
| All pages | 120 | 86 (71.7%) | 63 (52.5%) | 33 (27.5%) | About 15.4 days |
Illustrative composite data for an invented brand. Not real measurements.
For the third-party half-life, persistence crossed 50% between day 7 (66.7%) and day 14 (45.6%):
- Total drop: 66.7 − 45.6 = 21.1 points over 7 days
- Drop needed: 66.7 − 50 = 16.7 points
- Days needed: (16.7 ÷ 21.1) × 7 = 5.5 days
- Half-life: 7 + 5.5 = about 12.5 days
In this invented example, Quillmoor's own pages held on much longer than third-party pages. Real data might show the opposite. Either way, the split tells you where the churn comes from, and that changes what you do about it.
Step 7: Add the brand mention rate, with its margin of error
Persistence tracks links. You also want to know how often the brand is named. On day 0, Quillmoor is named in 34 of the 100 answers.
34 ÷ 100 = 34% mention rate
With only 100 answers, that number has a wide margin of error. A standard way to estimate it is:
1.96 × √(0.34 × 0.66 ÷ 100) = 1.96 × 0.0474 = 0.093, or about ±9 points
So the true mention rate could be anywhere from about 25% to 43%. If next month's number is 40%, that is not proof of improvement. It is still inside the range.
How many runs you need before you trust a number
There is no perfect number. A quick rule for the margin of error on a percentage is about 1 divided by the square root of the number of answers. This works best when the rate is somewhere in the middle, between about 20% and 80%.
| Answers collected | Rough margin of error | What it can detect |
|---|---|---|
| 25 (for example, 5 prompts × 5 runs) | about ±20 points | Only very large differences |
| 100 (20 prompts × 5 runs) | about ±10 points | Large changes, like 20% to 45% |
| 400 (20 prompts × 20 runs) | about ±5 points | Moderate changes, like 30% to 40% |
Treat these as best-case numbers. Answers to the same prompt tend to resemble each other, so 20 runs of one prompt carry less information than 20 different prompts. The real margin is usually a bit wider than the table shows.
For a sense of scale, the SparkToro study's author concluded that to really know an AI's set of recommendations "you need to ask over and over again; usually at least 60-100X, then average these out."1 He also wrote that "visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric."1
A practical setup for most brands looks like this:
- At least 20 prompts per topic, so one odd prompt cannot swing the result.
- At least 5 runs per prompt per check, more for prompts you care about most.
- A fixed schedule, such as weekly, on the same day and at a similar time.
- Each platform tracked separately. AI Overviews and AI Mode share so few citations that combining them hides what is going on in each.5
- A noise-floor check (the same-day re-run from Step 3) each time you change your prompt set or platform.
How to report churn to a client without over-reading noise
The biggest risk with unstable data is telling a story that is not there. A client sees a drop from 38% to 31% and asks what went wrong. Often, nothing did. These habits help.
- Always show the range, not just the number. Write "34% (±9 points, 100 answers)" instead of "34%." It reminds everyone the number is an estimate.
- Set a "real change" rule before you look. For example: only call a change real if it is bigger than the margin of error and points the same way two checks in a row.
- Don't report rank position. The SparkToro data shows that list order almost never repeats.1 A brand that is "number 2" today may be "number 5" an hour later for no reason.
- Report persistence next to the noise floor. "71.7% of links survived 7 days, against a same-day noise floor of 90%" tells a clearer story than either number alone.
- Separate links from mentions. Because the meaning of answers is steadier than their sources,4 a client can lose a cited link while still being named in the answer. Those are different problems with different fixes.
For more on reading competitor comparisons, see our guide on how to read a share-of-voice chart.
Here is the kind of wording that works well in a client report:
"Quillmoor was named in 34% of 100 ChatGPT answers this month, up from 29% last month. Both numbers have a margin of about 9 points, so this is not yet a confirmed increase. We will call it confirmed if next month's number is above 38%."
Example wording using the invented brand above.
What content actions plausibly help
We have not found a controlled test showing that a specific content change makes AI citations stickier. What follows rests on correlation studies or on how the systems are described to work. So these are reasonable bets, not proven fixes.
Keep important pages genuinely updated
Ahrefs analyzed about 17 million citations across AI assistants and organic Google results. The average URL cited by AI assistants was 1,064 days old, compared with 1,432 days for organic search results. That makes AI-cited content about 25.7% "fresher."10
The effect differed a lot by platform. ChatGPT cited URLs about 458 days newer than organic results. Google's AI Overviews showed almost no preference, citing pages 16 days older on average.10
Three cautions come with this finding:
- It is correlation, not causation. Newer pages might be cited more because they are newer, or because they happen to cover topics people are currently searching.
- Old content still wins a lot. The average cited page was still about 2.9 years old.10
- Fake updates do not count. The Ahrefs authors point to Google's John Mueller warning against changing publish dates without real changes to the page.10
So the useful action is real maintenance: new data, corrected facts, current pricing and fresh examples on the pages that matter most to your buyers.
Keep your facts consistent everywhere they appear
Recall the Ahrefs finding: the meaning of AI Overviews stayed very steady even as about half the sources rotated.4 The AI seems to keep landing on the same conclusion while swapping the pages it cites.
This suggests a logical step. If the AI is going to rotate through different pages about you, every one of those pages should say the same correct things. Your pricing, features, locations and key claims should match across your site, review profiles, directory listings and partner pages. Then, whichever page gets picked this week, the answer about you is more likely to stay accurate. We have not found a study that tests this directly, so treat it as good practice, not a proven ranking factor.
Aim to be part of the steady core, not a lucky link
The SparkToro data showed that some brands appear in most answers even when the exact list changes every time.1 Those brands are part of a stable "core" of options the AI keeps returning to. Other brands appear only now and then.
Your measurement shows which group you are in. If you appear in a small share of answers, a plausible goal is wider coverage on the third-party sources the AI already cites for your prompts. Your persistence data (Step 6) shows which of those sources the AI keeps returning to. This is a reasoned bet, not a tested result.
Don't chase every drop
Finally, the cheapest action is often to do nothing. If a cited link disappears for one week, the research says there is a fair chance it comes back on its own. Semrush found that 43% of removed URLs returned later in the same month.3 Acting on every bounce wastes time and makes it harder to see what any one change actually did.
See how stable your own AI visibility really is
Single snapshots can make a brand look like it is winning or losing when it is really just noise. Ripplix runs your buyers' questions repeatedly across AI platforms, so you can see which citations and mentions actually hold over time and which ones come and go.
Get your free AI Visibility Report →
Sources:
- SparkToro: NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility (January 27, 2026)
- arXiv: Characterizing Web Search in The Age of Generative AI, Kirsten et al., Max Planck Institute for Software Systems, Ruhr University Bochum and UA Ruhr (October 13, 2025, revised May 31, 2026)
- Semrush: Exploring URL Volatility in Google's AI Overviews (November 20, 2024)
- Ahrefs: AI Overviews Change Every 2 Days (But Never Change Their Mind) (November 11, 2025)
- Ahrefs: Are AI Mode and AI Overviews Just Different Versions of the Same Answer? (730K Responses Studied) (December 15, 2025)
- Google Search Central: AI features and your website (last updated December 10, 2025)
- OpenAI Help Center: Searching the web with ChatGPT (accessed October 7, 2026)
- Thinking Machines Lab: Defeating Nondeterminism in LLM Inference, Horace He (September 10, 2025)
- arXiv: Non-Determinism of "Deterministic" LLM Settings, Atil et al. (August 6, 2024, revised April 2, 2025)
- Ahrefs: New Study: AI Assistants Prefer to Cite "Fresher" Content (17 Million Citations Analyzed) (July 28, 2025)


