Ask ChatGPT "best sunscreen brands in India" twice and you will usually get two different lists. Published research puts the odds of the same prompt returning the same brand list twice at less than 1 in 100 — and the same list in the same order at roughly 1 in 1,000. If a tool tells you that you "rank #3 in ChatGPT", the honest question is: on which of the hundred different answers it could have sampled?
What the published studies actually found
Three independent studies in 2026 measured how unstable AI answers are:
- SparkToro (January 2026) ran 2,961 repeated prompts with 600 volunteers and found the same prompt returned the identical brand list less than 1% of the time. Even a category's dominant brand appeared in only 85–97% of answers in narrow categories — and far less in broad ones.
- Kevin Indig's analysis of 815,000 prompt–page pairs (June 2026) found only 2.2% of citations persisted across three ChatGPT runs of the same prompt, with within-model variance of 10–34%.
- SISTRIX measured that ChatGPT replaces roughly 74% of its cited sources week over week.
None of this means AI answers are random. Brands with strong footprints appear most of the time; weak brands appear sometimes. The signal is a probability, not a position.
Why single-run trackers mislead
Most AI-visibility tools poll each prompt once a day and display the result as a rank. Two problems follow:
- Day-to-day movement is mostly noise. If your brand appears in ~60% of answers, a daily coin-flip run will show you "disappearing" two days a week — and reappearing, with an alert each time.
- Point estimates hide sample size. "You rank #3" from one run and from forty-five runs look identical on a dashboard, but one is an anecdote and the other is a measurement.
What honest measurement looks like
The statistically defensible way to report AI visibility is the way pollsters report elections:
- A rate, not a rank: "appeared in 62% of sampled answers", not "ranked #3".
- A margin of error: with 45 sampled answers, 62% really means "somewhere between 54% and 70%" at 95% confidence (a Wilson interval).
- Sample size on every number: any score without its n is marketing, not measurement.
- Change-alerts only beyond the band: movement inside the margin of error is weather, not news.
This is how Depra reports visibility: every score carries its sample size and its 95% range, and the language split (English vs Hinglish prompts) is computed the same way. When a number moves inside its band, we say so instead of sending you an alert.
What this means for your brand
The variance is not a reason to ignore AI search — it is a reason to measure it properly. The brands that appear in 80%+ of answers got there through consistent, quotable coverage on the sources AI engines actually read (review sites, comparison pages, Reddit, category listicles). Improving your appearance rate is a real, trackable outcome. Chasing a single-run "rank" is not.
