You track brand mentions in AI answers by running a fixed set of buyer prompts across each engine on a schedule, then recording four things per run: whether your brand appears, where it appears, how it is described, and who appears alongside you. The catch that most guides skip is that AI answers are non-deterministic, so one run of one prompt tells you almost nothing. You need enough runs to turn a coin-flip into a rate.
This matters now because buyers are asking engines the questions they used to type into Google. Gartner expects traditional search volume to fall 25% by 2026 (Gartner, 2024). If ChatGPT, Gemini, Perplexity and Google AI Overviews are shaping which brands get shortlisted, then AI visibility is a metric you manage, and anything you manage you first have to measure honestly.
| Metric | What it answers | How to record it |
|---|---|---|
| Mention rate | How often are we visible at all? | Runs where the brand appears, divided by total runs, per prompt per engine |
| Position | Are we top-of-mind or an afterthought? | Rank of first mention: 1st named, top 3, or lower |
| Sentiment | Are we recommended or caveated? | Positive, neutral, or negative framing of the mention |
| Share of voice | Do engines treat us as a leader? | Our mention rate set against named competitors on identical prompts |
What should you measure when tracking brand mentions in AI answers?
Measure four things on every prompt run: mention rate, position, sentiment, and share of voice against competitors. Each answers a different business question, and tracking only one of them hides the others. A brand can be mentioned often but always last, or mentioned warmly but rarely, and you only see that gap when all four sit side by side.
Record these as numbers you can trend month over month, not as a vibe from reading a few answers. The point of the schema below is that any two people on your team score the same answer the same way.
- Mention rate (visibility) The share of runs for a prompt where your brand appears at all. If your brand shows up in 6 of 20 runs of "best CRM for small teams", that is a 30% mention rate, and that single number is the backbone of AI visibility.
- Position Where you appear inside the answer: first named, in the top three, or buried at the bottom. Being mentioned tenth in a list a buyer skims is close to being invisible, so position weights the raw mention.
- Sentiment How the engine frames you: recommended, listed neutrally, or flagged with a caveat like price or thin support. Sentiment is where reputation problems surface, because an engine can mention you often while quietly steering buyers elsewhere.
- Share of voice Your mention rate set against named competitors on the same prompts. This is the number leadership cares about, because it says whether ChatGPT and Gemini treat you as a category leader or an also-ran.
Why do single runs lie about your AI visibility?
Single runs lie because large language models are non-deterministic: ask ChatGPT or Gemini the same question twice and you can get two different brand lists. The models sample from a probability distribution over possible next words, so the same prompt produces different answers across runs, sessions, and dates. One flattering answer is not evidence you are winning, and one bad answer is not evidence you are losing.
Treat each prompt like a poll, not a fact lookup. If you asked five people on the street to name a good running shoe, you would not publish the first reply as the market view. The same discipline applies here. A brand that appears in 2 of 20 runs and a brand that appears in 18 of 20 look identical if you only ever run each prompt once.
As a working rule, run each prompt at least 10 to 20 times per engine per cycle, in fresh sessions with memory and personalisation off. Fewer runs and your mention rate swings on noise. More runs tightens the estimate but costs time, which is exactly the trade-off that pushes teams toward automation later. Also vary nothing else: same wording, same engine settings, so the only thing moving is the model's own sampling.
Gartner expects traditional search volume to fall 25% by 2026, which is why a trustworthy AI visibility number is becoming a board-level metric.(Gartner, 2024)
How do you build a prompt set that reflects real buyers?
Build your prompt set from questions real buyers actually ask, mined from Reddit, Quora, marketplace reviews and support tickets, rather than the polished phrases your marketing team would write. Buyers do not search "leading omnichannel commerce platform". They ask "cheapest way to sell on Instagram and my own site" or "is X worth it for a small kitchen". Track the second kind, because that is what engines are answering.
Anchor the set to buying intent, not brand vanity. Prompts that name your brand only tell you how engines describe you when asked. The prompts that decide revenue are category prompts where your name has to earn its place: "best", "cheapest", "alternative to a competitor", "good for a specific use case". Keep the wording fixed once chosen, because changing a prompt resets its trend line.
Good prompt sets also model your content well. Structured, well-cited pages earn up to 40% more AI citations (Princeton GEO study, KDD 2024), so the same buyer questions you track are the ones your content should answer clearly on-site.
- Mine, do not guess Pull real phrasing from Reddit threads, Quora questions, and Amazon or app-store reviews in your category. The awkward, specific wording is the signal.
- Cover the buying journey Include category discovery ("best tools for X"), comparison ("A vs B"), and objection prompts ("is X worth the price"). Each stage exposes a different weakness.
- Segment by use case and geography "Best CRM" and "best CRM for a 3-person agency in India" return different brands. Track the segments where you actually compete.
- Freeze the wording Lock each prompt once it is in the set. If you must reword one, treat it as a new prompt with a fresh baseline, so your trend stays honest.
What is a simple monthly method a team can follow?
A team can track AI visibility with a spreadsheet and a fixed monthly routine. The method below needs no engineering. It needs discipline: same prompts, same run counts, same scoring rules, on the same week each month. Consistency is what makes month two comparable to month one.
- Step 1: Fix your prompt set Choose 15 to 30 buyer prompts mined from real chatter. List them in a sheet with a stable ID per prompt so results line up across months.
- Step 2: Choose your engines Pick the engines your buyers use, for example ChatGPT, Gemini, Perplexity and Google AI Overviews. Test in fresh sessions with personalisation and chat memory switched off.
- Step 3: Run each prompt 10 to 20 times per engine Paste each prompt, record the answer, repeat. Yes, it is repetitive. That repetition is the measurement, because it converts non-determinism into a rate you can trust.
- Step 4: Score every run For each run log four columns: mentioned yes or no, position (1st, top 3, or lower), sentiment (positive, neutral, negative), and which competitors appeared. Agree the rules up front so two people score identically.
- Step 5: Roll up the metrics Per prompt and per engine, compute mention rate, average position, sentiment split, and share of voice versus named competitors. These four roll-ups are your scorecard.
- Step 6: Repeat on the same week monthly Re-run the identical set next month and compare. Rising mention rate and improving position mean your AEO and GEO work is landing. Flat or falling means revisit content and sources.
When should you use a tool instead of a spreadsheet?
Switch to a tool when the manual maths stops being feasible: many prompts, times several engines, times 10 to 20 runs, times monthly, quickly runs into thousands of answers to collect and score by hand. At that scale a spreadsheet breaks on human time and consistency, and scoring drift creeps in as different people read answers differently.
Platforms exist to automate exactly this loop. Tools like Otterly and Peec track prompts across engines and chart mention rate and share of voice over time. Depra AI runs the same fixed-prompt, multi-run method across seven engines, adds SKU-level product tracking for D2C catalogues, and mines prompts from real buyer chatter rather than guesses. Each is a fair choice depending on whether you want self-serve software or a service that ships the fixes too.
The honest version: start manual for one quarter if you have never measured this. It teaches you what the metrics mean and which prompts matter. Once you know the routine is worth keeping, a tool removes the tedium and the human error, and frees your team to act on the report instead of assembling it.
Structured, well-cited pages earn up to 40% more AI citations, so tracking is only half the job: the other half is fixing the pages engines cite.(Princeton GEO study, KDD 2024)