All resources
Blog

AI answers change every time you ask. How many checks do you need before you trust the number?

Studies from SparkToro, Detailed and Ahrefs show AI answers shift from run to run. Here is a plain-English sample-size method, with the math worked out, for knowing when a change in your AI mention rate is real.

Ask ChatGPT the same question twice and you will often get two different answers. Different brands show up. The order changes. The sources it links to change too. So if you checked once last Tuesday and saw your brand mentioned, that single check tells you much less than it seems.

This piece looks at what the research says about how much AI answers move around. Then it gives you a simple method, with the math worked out, for deciding how many checks you need before you can trust the number.

What the research shows about AI answers changing

Several independent teams have now measured this. They used different methods, different platforms and different time periods. They all found the same basic thing: AI answers are unstable from one run to the next.

SparkToro: almost every answer is unique

In January 2026, Rand Fishkin of SparkToro, an audience research company, published a study on how consistent AI tools are when they recommend brands.1 Here is how it worked:

  • Volunteers ran 12 different prompts through ChatGPT, Claude and Google's AI answers (AI Overviews, or AI Mode when no AI Overview showed). The runs took place in November and December 2025.1
  • In SparkToro's words, "600 volunteers ran 12 different prompts through each of the 3 tools a combined 2,961 times."1
  • Each prompt was run 60 to 100 times per platform, according to Search Engine Journal's coverage.2
  • The prompts asked for recommendations in areas like chef's knives, headphones, cancer care hospitals, digital marketing consultants and science fiction novels.2

The main result was blunt. SparkToro wrote: "there's a <1 in 100 chance that ChatGPT or Google's AI, if asked 100X, will give you the same list of brands in any two responses."1 Getting the same list in the same order was rarer still, under 1 in 1,000 times.1

The study also tested something closer to real life. Real people don't type the same prompt. So SparkToro ran 142 different human-written prompts about choosing headphones for a family member who travels, each several times. That produced 994 responses.1 The prompts were very different from each other. Search Engine Journal reported that "The semantic similarity score across those human-written prompts was 0.081."2 (A semantic similarity score measures how close two pieces of text are in meaning. A score near 1 means nearly the same; 0.081 is very low.)

Even so, a few brands kept appearing. SparkToro found that "headphones like Bose, Sony, Sennheiser, and Apple showed up 55-77% of the time."1

This is the key point for anyone tracking AI visibility. Any single answer is close to random. But across many answers, a pattern appears. SparkToro's conclusion was that "Measuring your brand's presence in AI answers with precision is a fool's errand." In the same section, Fishkin wrote that he now believes "visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric."1

He also left an open question that this article tries to answer in practical terms: "How many times do you need to run a prompt to have statistically sound answers about a brand's relative visibility?"1

Detailed: the top brands are steady, the rest are not

In September 2026, Glen Allsopp of the SEO publisher Detailed tracked AI answers every day for four weeks.3 In his words: "I've looked at over 70,000 responses, checking 1,300+ prompts daily in both ChatGPT and AI Overviews, to see how they shift over four weeks."3 Each prompt was checked once a day from the US, using a logged-out account, from August 26 to September 22, 2026.3

His results add an important detail. Not every brand in an answer moves around the same amount.

Finding (Detailed, 28 days)ChatGPTGoogle AI Overviews
Share of days the leading brand appeared (typical prompt)82%89%
Two random days with the exact same set of brands (typical prompt)0.3% of pairs1.1% of pairs
Cited pages that were cited again the next dayAbout a quarterAround half

Source: Detailed.3

Allsopp split brands into two groups. "Core brands" were named on at least 80% of days. The "tail" was everything else.3 He found that "On ChatGPT, the core brands changed by 13% from one day to the next, against 78% for the tail."3

Two more findings matter for tracking:

  • "41% of ChatGPT prompts never returned the same set of brands twice, not even on one pair of days in the month."3
  • "For the same prompt, the two platforms agreed on the most-present brand only 27% of the time."3

So the brand that leads on ChatGPT is often not the brand that leads on AI Overviews. That's why you should never blend the two platforms into one number and compare it with last month's blended number.

Ahrefs: AI Overviews change every couple of days

Ahrefs, an SEO software company, studied Google's AI Overviews (the AI summaries at the top of some Google results). Its team looked at more than 43,000 keywords over one month. Each keyword had at least 16 recorded AI Overviews.4

  • "AI Overviews have a 70% chance of changing from one observation to the next."4
  • On average, the content of an AI Overview tended to change every 2.15 days.4
  • Between one response and the next, only 54.5% of the cited web addresses (URLs) stayed the same on average. That means 45.5% of cited sources were new each time.4

There is an interesting twist. The meaning of the answers stayed very similar. Ahrefs measured an average cosine similarity score of 0.95 between consecutive AI Overviews, "where 1.0 represents a perfect match."4 (Cosine similarity is another way to measure how close two texts are in meaning.)

In other words, the overall message stays about the same, but the exact wording, brands and sources keep shifting. For a brand, those details are the part that matters. Being named, or being the linked source, is exactly what changes.

Why AI answers change even when nothing else has

It's tempting to assume a changed answer means something happened. Maybe a competitor published new content, or Google updated something. Sometimes that's true. But a lot of the change happens with no outside cause at all.

AI chat tools build answers one word at a time. At each step, the model picks from several likely next words. A setting called "temperature" controls how much randomness goes into that pick. Higher temperature means more variety.

You might think setting temperature to zero would make answers identical every time. Research says it does not fully do that:

  • A 2024 paper on arXiv tested five large language models set up to be "deterministic," meaning they should give the same output every time. The authors ran eight common tasks 10 times each. They saw "accuracy variations up to 15% across naturally occurring runs."5 They also wrote that "none of the LLMs consistently delivers repeatable accuracy across all tasks, much less identical output strings."5
  • In September 2025, Horace He at Thinking Machines Lab asked one open model (Qwen3-235B) to "Tell me about Richard Feynman" 1,000 times at temperature zero. It produced 80 different completions.6 The first difference showed up at word-piece (token) 103: 992 versions said "Queens, New York" and 8 said "New York City."6 Only after the team changed how the model's math was run on the hardware did all 1,000 answers come out the same.6

The consumer apps your customers use add even more variety. Answers can depend on the exact wording of the question, the location, whether the person is logged in, and live web results the tool pulls in. SparkToro's headphone test showed how different real people's wording is.1

So treat each AI answer as one draw from a range of possible answers. One draw can't tell you the shape of the range. You need many draws.

How many checks you need: the method in plain words

This section is Ripplix's practical answer to SparkToro's open question. It uses a standard formula from statistics. The same formula is used to work out the "margin of error" you see in opinion polls.

The formula, one step at a time

Start with two numbers:

  • p is your brand's true mention rate. That's how often you'd be mentioned if you could run a prompt an endless number of times. For example, 0.3 means 30% of answers.
  • n is the number of answers you actually checked.

The margin of error is roughly:

margin = 1.96 × √( p × (1 − p) ÷ n )

Here is what each part does:

  1. p × (1 − p) measures how mixed the results are. It's largest at 50%, when "mentioned" and "not mentioned" are equally likely.
  2. Dividing by n shrinks the error as you check more answers.
  3. The square root (√) is why more checks help less and less. To cut the error in half, you need four times as many checks, not twice as many.
  4. 1.96 sets the confidence level at 95%. Roughly, if you repeated your whole check many times, about 95 out of 100 of the ranges you got would contain the true rate.

A worked example: say your true rate is 30% and you check 10 answers.

  • p × (1 − p) = 0.3 × 0.7 = 0.21
  • 0.21 ÷ 10 = 0.021
  • √0.021 = about 0.145
  • 1.96 × 0.145 = about 0.284, or ±28.4 points

So with 10 checks, a measured 30% really means "somewhere from about 2% to about 58%." That range is too wide to make decisions with.

The margin of error at different numbers of checks

Number of answers checked (n)Margin if true rate is 30%Likely rangeMargin if true rate is 50%Likely range
5±40.2 points0% to 70%±43.8 points6% to 94%
10±28.4 points2% to 58%±31.0 points19% to 81%
20±20.1 points10% to 50%±21.9 points28% to 72%
50±12.7 points17% to 43%±13.9 points36% to 64%
100±9.0 points21% to 39%±9.8 points40% to 60%
200±6.4 points24% to 36%±6.9 points43% to 57%

Ripplix calculation using the formula above. Ranges are rounded to the nearest whole point. At 5 checks the formula's lower bound for 30% falls below zero, so it is shown as 0%.

Two quick rules of thumb come out of this table:

  • To get within about ±10 points, you need around 100 answers (97 at a 50% rate).
  • To get within about ±5 points, you need around 385 answers at a 50% rate.

That sounds like a lot. But "answers" doesn't have to mean one prompt run 100 times. It can be 25 prompts run 4 times each. That also matches SparkToro's advice to measure "across dozens to hundreds of prompts run multiple times."1 We come back to this below.

Why a jump from 30% to 40% on 10 runs is not a real change

Say you ran a prompt 10 times last month and your brand showed up 3 times (30%). This month you ran it 10 times again and showed up 4 times (40%). It looks like a 10-point gain. It almost certainly isn't one you can prove.

Here is why, in three steps:

  1. The ranges overlap a lot. From the table, 30% on 10 runs means roughly 2% to 58%. A 40% result on 10 runs has a similarly wide range. The two ranges almost completely cover each other.
  2. The gap needed is bigger than the gap you saw. When you compare two measurements, both carry error. Using the same formula for the difference between 30% and 40%, each on 10 runs, the margin on the gap is about ±41.6 points. Your gap was 10 points. That's well inside the noise.
  3. Chance alone does this often. If your true rate never moved from 30%, the chance of seeing 4 or more mentions out of 10 is about 35%, roughly 1 in 3. So a "gain" like this would show up about one month in three, even if nothing had changed at all.

How many checks would you need to trust a 30% to 40% move? Using the same formula, about 173 answers in each period before a 10-point gap clears the margin. With 100 answers per period the margin on the gap is still about ±13.1 points. With 200 it drops to about ±9.3 points.

A single yes or no tells you almost nothing. A mention rate from 10 runs tells you a little. A rate from 100 or more answers, compared with the same setup over time, is what you can actually act on.

The honest limits of this math

The formula is a good rough guide, not a perfect one. Keep these limits in mind:

  • It assumes each run is independent. Runs done in the same minute, the same session or the same location may be more alike than truly separate runs. If so, the real margin is wider than the table shows. Spreading runs over different days helps.
  • It's weaker at small numbers and at extremes. With very few checks, or rates near 0% or 100%, the simple formula gets less accurate. That's why the 5-run row hits the 0% floor. Treat those rows as "very uncertain," not as exact ranges.
  • It only covers run-to-run noise. It can't fix a bad prompt list. If your prompts don't match what buyers actually ask, a precise number is still precisely wrong. Our guide to finding the questions buyers really ask covers how to build that list.

An illustrative example: one brand, eight weeks

This is an illustrative composite, not a real client or a real study. The brand and the numbers are invented to show how the method works in practice.

"Fernway" is a made-up project management tool. Its team tracks 20 buyer prompts on ChatGPT. They run each prompt 3 times a week, always logged out, from the same country. That gives them 60 answers per week. They count how many answers mention Fernway.

WeekAnswers mentioning Fernway (of 60)Mention rateMargin of error
11830.0%±11.6 points
22440.0%±12.4 points
31931.7%±11.8 points
42135.0%±12.1 points
52541.7%±12.5 points
62745.0%±12.6 points
72643.3%±12.5 points
82948.3%±12.6 points

Illustrative numbers. Margins calculated with the formula above.

Here is how the team should read this:

  • Week 1 to week 2 is not news. It moved from 30% to 40%. But the margin on the gap between two 60-answer weeks is about ±17.0 points. A 10-point jump is inside that. Then week 3 fell back to 31.7%, which shows how much a single week bounces.
  • Grouping weeks gives a clearer picture. Weeks 1 to 4 add up to 82 mentions out of 240 answers, or 34.2% (±6.0 points). Weeks 5 to 8 add up to 107 out of 240, or 44.6% (±6.3 points).
  • Now the change is likely real. The gap between the two four-week blocks is 10.4 points. The margin on that gap is about ±8.7 points. The gap is bigger than the margin, so it's unlikely to be just noise.

Notice what this does and doesn't tell the Fernway team. It tells them their mention rate on ChatGPT very likely went up. It does not tell them why. If they published new comparison pages in week 4, that's a reasonable lead to look into. But the timing alone is a correlation, not proof that the pages caused the rise. Competitors, model updates and changes in the live web results could all play a part.

A practical tracking routine you can run

Here's a routine built on the research and the math above. You can run it by hand with a spreadsheet, or with a tool.

  1. Fix your prompt set. Pick a set of real buyer questions and keep it the same from week to week. If you change the prompts, you can't compare the numbers. Add new prompts as a separate group instead of mixing them in.
  2. Run each prompt several times per platform. Three to five runs per prompt is a sensible start. With 20 to 30 prompts, that gets you to roughly 60 to 150 answers per platform per period. Spread the runs over different days rather than firing them all at once.
  3. Keep the conditions the same. Same platform, same country or city, same logged-out (or same logged-in) state, same type of account. Detailed found ChatGPT and AI Overviews agreed on the top brand for a prompt only 27% of the time.3 So each platform needs its own number.
  4. Report a rate with a range. Write "mentioned in 38% of answers (±9 points)" instead of "yes, we're in ChatGPT." A rate with a range is honest about what you know.
  5. Separate your steady results from your shaky ones. Detailed showed that core brands are far more stable than tail brands.3 If you appear in almost every answer for a prompt, small dips don't matter much. If you appear only some of the time, expect big swings and judge on longer windows.
  6. Watch trends over weeks, not single checks. Compare four-week blocks, or use a rolling average. Only call something a change when the gap is bigger than the margin.
  7. Track sources separately from mentions. Cited pages change faster than brand mentions. Ahrefs found that only 54.5% of cited URLs carried over between consecutive AI Overviews on average.4 A page that drops out of the citations once hasn't necessarily "lost" anything yet.
  8. Don't track "rank position" in AI answers. The order of brands in a list is the least stable part of an answer. SparkToro found the same list in the same order came up less than 1 in 1,000 times.1 Mention rate and share of voice are more useful. Our guide on how to read share of voice explains how to compare your rate with competitors'.

Quick reference: what a result is worth

What you checkedWhat you can honestly say
One run of one prompt"We appeared once." Nothing about how often.
10 runsA very rough rate, roughly ±30 points. Fine for spotting "almost never" versus "almost always."
About 50 answersA usable rough rate, roughly ±13 points. Not enough to trust small moves.
About 100 answersA solid rate, roughly ±10 points.
About 175 or more answers per period, same setupEnough to start trusting a 10-point change between two periods.

Ripplix calculation, based on mention rates between 30% and 50%.

What this means for how you report AI visibility

The research points in one direction. Single answers are unstable, and the instability is built into how these tools work.56 Brand lists almost never repeat exactly.1 Cited sources turn over quickly.34 But patterns across many answers are real and measurable. Leading brands showed up on most days in Detailed's data, and a few headphone brands kept appearing across very different prompts in SparkToro's test.13

So the useful question isn't "are we in ChatGPT?" It's "in what share of answers do we appear, how sure are we of that number, and is it moving?"

For a team reporting to leadership, this changes the conversation. A screenshot of one answer is a story, not a measurement. A mention rate with a range, tracked the same way over weeks, is something leadership can actually plan around.

Move from one screenshot to a real baseline

Ripplix tracks brand mentions, competitors and citations across AI answers on platforms like ChatGPT, Google AI Overviews, Gemini, Claude and Perplexity. The free AI Visibility Report is a quick way to see where your brand stands today, so you have a starting point to measure the trend against.

Get your free AI Visibility Report →

Sources:

  1. SparkToro: NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility (January 27, 2026)
  2. Search Engine Journal: AI Recommendations Change With Nearly Every Query: Sparktoro (January 30, 2026)
  3. Detailed: I Tracked 70K+ ChatGPT and AI Overview Responses to Measure Their Volatility (September 23, 2026)
  4. Ahrefs: AI Overviews Change Every 2 Days (But Never Change Their Mind) (November 11, 2025, updated August 12, 2026)
  5. arXiv: Non-Determinism of "Deterministic" LLM Settings, Atil et al. (August 2024, revised April 2025)
  6. Thinking Machines Lab: Defeating Nondeterminism in LLM Inference, by Horace He (September 10, 2025)