Measuring AI visibility means tracking the rate at which your brand gets named in answers to a fixed set of buying prompts, sampled repeatedly, reported separately for each engine, with the sample size and a confidence interval on every number. One screenshot of a good answer is not a measurement, and screenshots tell you whatever you want to hear.
What follows: why one run proves nothing, why appearance rate beats "rank", how to read a confidence band without a statistics degree, why a blended cross-engine score is worse than no score, and the five metrics that survive scrutiny.
A single run is an anecdote, not a measurement
The same prompt does not return the same answer twice. SparkToro logged 2,961 runs from 600 volunteers in January 2026 and found the odds of one prompt returning the identical brand list twice are under 1 in 100. The same list in the same order is around 1 in 1,000.
Citations are less stable still. Kevin Indig's June 2026 analysis of 815,000 prompt-page pairs found only 2.2% of citations persisted across three consecutive ChatGPT runs, with within-model variance of 10% to 34%. SISTRIX has measured ChatGPT replacing roughly 74% of its cited sources week over week.
So a screenshot of your brand in an answer is not proof visibility improved, and one naming your competitor instead is not proof it dropped. You are measuring a coin of unknown bias: one flip tells you nothing, enough flips tell you a lot.
Measure appearance rate, not "rank"
There is no position 1 in an AI answer. The output is a paragraph or short list of brands that changes shape every run, so keyword-rank thinking produces a number you cannot defend.
The primary metric is appearance rate: the share of runs, for a given prompt set and engine, in which your brand is named at all. Run 40 prompts on ChatGPT three times each, get named in 74 of those 120 answers, and your appearance rate is 62%.
Position is secondary. Track average position only across answers where you appear, and never report it without the appearance rate beside it. Being named first in 8% of answers is worse than being named third in 55% of them, and a lone position number is the most common way these dashboards flatter their owners.
Sample size and the confidence band, in plain language
Every rate needs two companions: n, the number of observations behind it, and a confidence interval, the range the true rate plausibly sits in. Here n is prompts multiplied by engines multiplied by runs. A dashboard that reports 62% but will not say what n is has given you nothing.
A worked example. Your brand appears in 28 of 45 observations, so 62%. The 95% Wilson interval on that runs roughly 48% to 75%, and the honest reading is "somewhere between about half and three quarters of the time". A follow-up month at 68% is not an improvement. It sits well inside the same band.
Narrowing the band takes more observations, not more analysis. At the same 62% rate, n of about 140 pulls the interval to roughly 54% to 70%, or plus or minus 8 percentage points. Two rules follow:
- Report the band, not just the point. "62% (n=45, 95% CI 48-75%)" is defensible. "62%" is not.
- Alert only on moves outside the band. If the new reading overlaps the old interval, nothing has been detected. Depra ships n and a 95% Wilson interval with every score, and raises a change alert only when a move clears the band.
One blended score hides the answer
The case against a single "AI visibility score" is empirical. Kevin Indig and Omnia analysed 3.7 million citations across 20,000 prompts in May 2026 and found only 2.37% of cited URLs appeared in all three of ChatGPT, Perplexity and Google AI Overviews. 91% appeared in exactly one engine. Peec's March 2026 study of 30 million sources found per-engine source preferences differ substantially, with Reddit the most-cited domain overall.
These engines are not four windows onto one index. They are four different selection processes, so a blended score averages away your only actionable information. A brand at 70% on Gemini and 10% on Perplexity reports the same 40% as a brand flat at 40% everywhere, and the two need completely different work. Report per engine, always. AEO vs GEO vs SEO covers why they diverge; Depra scores ChatGPT, Google AI Overviews, Gemini and Perplexity separately for the same reason.
What a defensible setup looks like
Miss any one of these and the trend line stops meaning anything.
- A frozen prompt set. Twenty to fifty prompts in real buying language, written once and left alone. Editing prompts mid-quarter silently resets your baseline. Add new ones as a separate cohort.
- Repeated sampling on a schedule. Daily where the engine allows it, weekly at minimum, and never on demand, because you will run it when you expect good news. Depra samples daily, with Perplexity weekly on paid plans.
- A rolling window. Report a trailing 7 or 28 days rather than the latest single pull, so day-level noise averages out.
- Alerting only beyond the interval. Set the trigger at the confidence band. Everything inside it is variance wearing a costume.
- Version stamping. Mark the date a provider ships a new model. A step change on that day is the engine changing, not your content working.
- Competitors on identical prompts. Relative share is far more stable than absolute rate, because a model update that moves everyone moves everyone.
Indian D2C brands need a seventh: track Hinglish alongside English and keep the split visible. India is ChatGPT's second-largest market at roughly 100 million weekly users, per Sam Altman in February 2026, and Google AI Mode launched here in June 2025, adding Hindi in September 2025. Real buyers type "best sunscreen for oily skin under 500 rupees" in Hinglish, not textbook Devanagari Hindi, and a brand can be visible in English while absent in Hinglish for the same category.
The five metrics worth tracking
| Metric | What it answers | How to read it honestly |
|---|---|---|
| Appearance rate | How often are we named at all? | Always with n and a confidence interval, split by engine and language |
| Share of voice | Of all brand mentions on our prompt set, what share is ours? | Against the same competitors on identical prompts, not a category average |
| Average position | When named, how prominently? | Only across answers where you appear, never without appearance rate |
| Sentiment | Are we recommended, hedged, or named as the thing to avoid? | Watch direction over weeks; single-answer sentiment is noise |
| Cited sources | Which domains does the engine pull from for our category? | Compare your citation footprint to competitors' to find the gap to close |
Cited sources is what turns measurement into a to-do list. Ahrefs' study of 75,000 brands found branded web mentions correlate with AI citations at rho=0.664, far above backlinks at rho=0.218, so the domains an engine keeps citing for your category map fairly directly onto where you need to exist. If you are absent from all of them, why ChatGPT doesn't mention your brand walks the failure patterns.
Three ways marketers fool themselves
Assuming a popular tactic worked. The Princeton GEO paper at KDD 2024 found that adding statistics, quotes and citations raised visibility by roughly 30-40%, but in a sandbox running GPT-3.5. C-SEO Bench at NeurIPS 2025 found many conversational-SEO tactics largely ineffective and sometimes negative, and Ahrefs reported in 2026 that 97% of llms.txt files received zero bot requests. Treat every tactic as a hypothesis and test it against your own prompt set.
Watching only your own site. The Ahrefs correlation points at third-party mentions, not on-page tweaks, so measurement scoped to your domain will keep showing you nothing.
Measuring after the fact. A prompt set built after a campaign launched has already absorbed the effect you wanted to detect. Baseline first.
Where tooling helps
A spreadsheet handles one engine and twenty prompts. It stops being viable at four engines, two languages, daily sampling and competitor benchmarking, which runs to thousands of observations a month. Depra's free plan covers 5 prompts weekly across ChatGPT, Gemini and AI Overviews, with paid tiers on INR pricing, and the D2C tool roundup compares the alternatives. Whatever you use, the buying criteria come down to four questions: does it show sample size, does it publish a confidence interval, does it report per engine, and does it run your competitors on identical prompts. A tool that fails any of those is selling you screenshots with a dashboard around them.
