Method

How to Measure AI Visibility (Without Fooling Yourself)

The honest method for measuring AI visibility: appearance rate over a fixed prompt set, sampled repeatedly, reported per engine, with sample size and a confidence interval on every number.

14 Aug 2026 · 7 min read

A horizontal axis with a shaded confidence band around a central point estimate, illustrating a margin of error rather than a single number

Measuring AI visibility means tracking the rate at which your brand gets named in answers to a fixed set of buying prompts, sampled repeatedly, reported separately for each engine, with the sample size and a confidence interval on every number. One screenshot of a good answer is not a measurement, and screenshots tell you whatever you want to hear.

What follows: why one run proves nothing, why appearance rate beats "rank", how to read a confidence band without a statistics degree, why a blended cross-engine score is worse than no score, and the five metrics that survive scrutiny.

A single run is an anecdote, not a measurement

The same prompt does not return the same answer twice. SparkToro logged 2,961 runs from 600 volunteers in January 2026 and found the odds of one prompt returning the identical brand list twice are under 1 in 100. The same list in the same order is around 1 in 1,000.

Citations are less stable still. Kevin Indig's June 2026 analysis of 815,000 prompt-page pairs found only 2.2% of citations persisted across three consecutive ChatGPT runs, with within-model variance of 10% to 34%. SISTRIX has measured ChatGPT replacing roughly 74% of its cited sources week over week.

So a screenshot of your brand in an answer is not proof visibility improved, and one naming your competitor instead is not proof it dropped. You are measuring a coin of unknown bias: one flip tells you nothing, enough flips tell you a lot.

Measure appearance rate, not "rank"

There is no position 1 in an AI answer. The output is a paragraph or short list of brands that changes shape every run, so keyword-rank thinking produces a number you cannot defend.

The primary metric is appearance rate: the share of runs, for a given prompt set and engine, in which your brand is named at all. Run 40 prompts on ChatGPT three times each, get named in 74 of those 120 answers, and your appearance rate is 62%.

Position is secondary. Track average position only across answers where you appear, and never report it without the appearance rate beside it. Being named first in 8% of answers is worse than being named third in 55% of them, and a lone position number is the most common way these dashboards flatter their owners.

Sample size and the confidence band, in plain language

Every rate needs two companions: n, the number of observations behind it, and a confidence interval, the range the true rate plausibly sits in. Here n is prompts multiplied by engines multiplied by runs. A dashboard that reports 62% but will not say what n is has given you nothing.

A worked example. Your brand appears in 28 of 45 observations, so 62%. The 95% Wilson interval on that runs roughly 48% to 75%, and the honest reading is "somewhere between about half and three quarters of the time". A follow-up month at 68% is not an improvement. It sits well inside the same band.

Narrowing the band takes more observations, not more analysis. At the same 62% rate, n of about 140 pulls the interval to roughly 54% to 70%, or plus or minus 8 percentage points. Two rules follow:

  • Report the band, not just the point. "62% (n=45, 95% CI 48-75%)" is defensible. "62%" is not.
  • Alert only on moves outside the band. If the new reading overlaps the old interval, nothing has been detected. Depra ships n and a 95% Wilson interval with every score, and raises a change alert only when a move clears the band.

One blended score hides the answer

The case against a single "AI visibility score" is empirical. Kevin Indig and Omnia analysed 3.7 million citations across 20,000 prompts in May 2026 and found only 2.37% of cited URLs appeared in all three of ChatGPT, Perplexity and Google AI Overviews. 91% appeared in exactly one engine. Peec's March 2026 study of 30 million sources found per-engine source preferences differ substantially, with Reddit the most-cited domain overall.

These engines are not four windows onto one index. They are four different selection processes, so a blended score averages away your only actionable information. A brand at 70% on Gemini and 10% on Perplexity reports the same 40% as a brand flat at 40% everywhere, and the two need completely different work. Report per engine, always. AEO vs GEO vs SEO covers why they diverge; Depra scores ChatGPT, Google AI Overviews, Gemini and Perplexity separately for the same reason.

What a defensible setup looks like

Miss any one of these and the trend line stops meaning anything.

  • A frozen prompt set. Twenty to fifty prompts in real buying language, written once and left alone. Editing prompts mid-quarter silently resets your baseline. Add new ones as a separate cohort.
  • Repeated sampling on a schedule. Daily where the engine allows it, weekly at minimum, and never on demand, because you will run it when you expect good news. Depra samples daily, with Perplexity weekly on paid plans.
  • A rolling window. Report a trailing 7 or 28 days rather than the latest single pull, so day-level noise averages out.
  • Alerting only beyond the interval. Set the trigger at the confidence band. Everything inside it is variance wearing a costume.
  • Version stamping. Mark the date a provider ships a new model. A step change on that day is the engine changing, not your content working.
  • Competitors on identical prompts. Relative share is far more stable than absolute rate, because a model update that moves everyone moves everyone.

Indian D2C brands need a seventh: track Hinglish alongside English and keep the split visible. India is ChatGPT's second-largest market at roughly 100 million weekly users, per Sam Altman in February 2026, and Google AI Mode launched here in June 2025, adding Hindi in September 2025. Real buyers type "best sunscreen for oily skin under 500 rupees" in Hinglish, not textbook Devanagari Hindi, and a brand can be visible in English while absent in Hinglish for the same category.

The five metrics worth tracking

MetricWhat it answersHow to read it honestly
Appearance rateHow often are we named at all?Always with n and a confidence interval, split by engine and language
Share of voiceOf all brand mentions on our prompt set, what share is ours?Against the same competitors on identical prompts, not a category average
Average positionWhen named, how prominently?Only across answers where you appear, never without appearance rate
SentimentAre we recommended, hedged, or named as the thing to avoid?Watch direction over weeks; single-answer sentiment is noise
Cited sourcesWhich domains does the engine pull from for our category?Compare your citation footprint to competitors' to find the gap to close

Cited sources is what turns measurement into a to-do list. Ahrefs' study of 75,000 brands found branded web mentions correlate with AI citations at rho=0.664, far above backlinks at rho=0.218, so the domains an engine keeps citing for your category map fairly directly onto where you need to exist. If you are absent from all of them, why ChatGPT doesn't mention your brand walks the failure patterns.

Three ways marketers fool themselves

Assuming a popular tactic worked. The Princeton GEO paper at KDD 2024 found that adding statistics, quotes and citations raised visibility by roughly 30-40%, but in a sandbox running GPT-3.5. C-SEO Bench at NeurIPS 2025 found many conversational-SEO tactics largely ineffective and sometimes negative, and Ahrefs reported in 2026 that 97% of llms.txt files received zero bot requests. Treat every tactic as a hypothesis and test it against your own prompt set.

Watching only your own site. The Ahrefs correlation points at third-party mentions, not on-page tweaks, so measurement scoped to your domain will keep showing you nothing.

Measuring after the fact. A prompt set built after a campaign launched has already absorbed the effect you wanted to detect. Baseline first.

Where tooling helps

A spreadsheet handles one engine and twenty prompts. It stops being viable at four engines, two languages, daily sampling and competitor benchmarking, which runs to thousands of observations a month. Depra's free plan covers 5 prompts weekly across ChatGPT, Gemini and AI Overviews, with paid tiers on INR pricing, and the D2C tool roundup compares the alternatives. Whatever you use, the buying criteria come down to four questions: does it show sample size, does it publish a confidence interval, does it report per engine, and does it run your competitors on identical prompts. A tool that fails any of those is selling you screenshots with a dashboard around them.

Frequently asked questions

How many times should I run the same prompt to measure AI visibility?

Enough runs that your confidence interval is narrow enough to act on. AI answers are non-deterministic: SparkToro's January 2026 study of 2,961 runs from 600 volunteers found the odds of one prompt returning the identical brand list twice are under 1 in 100. In practice, run 20 to 50 prompts at least weekly on each engine and let observations accumulate. Always report the appearance rate with its sample size and a 95% confidence interval, and treat any move inside that band as noise rather than a result.

What counts as a good AI visibility score or appearance rate?

There is no universal benchmark, because appearance rate depends entirely on how competitive your prompt set is. A 30% rate on a crowded category prompt can be stronger than 80% on a niche one. The useful comparison is relative: run the identical prompts for three or four named competitors and track your share against theirs over time. Relative share is also more stable than absolute rate, because a model update that moves everyone moves everyone.

How do you calculate share of voice in AI answers?

Count every brand mention returned across your prompt set for a given engine and time window, then divide your brand's mentions by that total. The number only means something if every brand was measured on identical prompts, in the same runs, over the same window. Report it per engine and alongside the raw appearance rate, because share of voice can rise simply because a competitor dropped out rather than because you gained ground.

Why did my AI visibility drop when I changed nothing?

Most likely variance or a model update, not your content. Kevin Indig's June 2026 analysis of 815,000 prompt-page pairs found only 2.2% of citations persisted across three consecutive ChatGPT runs, and SISTRIX has measured ChatGPT replacing roughly 74% of its cited sources week over week. Check whether the new reading falls outside the previous confidence interval. If the intervals overlap, nothing has been detected. If the drop lands on the date an engine shipped a new model, that is the engine changing, not you.

Can I use one blended AI visibility score across all engines?

No, and it will actively mislead you. Kevin Indig and Omnia found in May 2026 that across 3.7 million citations from 20,000 prompts, only 2.37% of cited URLs appeared in all three of ChatGPT, Perplexity and Google AI Overviews, while 91% appeared in exactly one engine. A brand at 70% on one engine and 10% on another averages to the same score as a brand sitting flat across both, and those two situations need completely different work. Report every engine separately.

Do I need to track Hinglish prompts separately from English?

Yes, if you sell to Indian consumers. Indians typically type buying queries in Hinglish rather than Devanagari Hindi, so a prompt set written only in English misses a large slice of real demand. India is ChatGPT's second-largest market at roughly 100 million weekly users, per Sam Altman in February 2026, and Google AI Mode launched in India in June 2025 and added Hindi in September 2025. Track both languages and keep the split visible, because a brand can be well covered in English and invisible in Hinglish for the same category.

See how AI engines talk about your brand

Track ChatGPT, Gemini, Perplexity & AI Overviews in English and Hinglish. First scan runs in minutes, free.

Run my free audit