Methodology

How to track brand mentions in AI answers

A repeatable method for measuring AI visibility across ChatGPT, Gemini, Perplexity and Google AI Overviews, built so single-run noise cannot fool you.

Updated 21 Jul 2026 · 9 min read

You track brand mentions in AI answers by running a fixed set of buyer prompts across each engine on a schedule, then recording four things per run: whether your brand appears, where it appears, how it is described, and who appears alongside you. The catch that most guides skip is that AI answers are non-deterministic, so one run of one prompt tells you almost nothing. You need enough runs to turn a coin-flip into a rate.

This matters now because buyers are asking engines the questions they used to type into Google. Gartner expects traditional search volume to fall 25% by 2026 (Gartner, 2024). If ChatGPT, Gemini, Perplexity and Google AI Overviews are shaping which brands get shortlisted, then AI visibility is a metric you manage, and anything you manage you first have to measure honestly.

A four-metric scoring schema for tracking brand mentions in AI answers
MetricWhat it answersHow to record it
Mention rateHow often are we visible at all?Runs where the brand appears, divided by total runs, per prompt per engine
PositionAre we top-of-mind or an afterthought?Rank of first mention: 1st named, top 3, or lower
SentimentAre we recommended or caveated?Positive, neutral, or negative framing of the mention
Share of voiceDo engines treat us as a leader?Our mention rate set against named competitors on identical prompts
Score every run against fixed rules so any two people on the team produce the same numbers.

What should you measure when tracking brand mentions in AI answers?

Measure four things on every prompt run: mention rate, position, sentiment, and share of voice against competitors. Each answers a different business question, and tracking only one of them hides the others. A brand can be mentioned often but always last, or mentioned warmly but rarely, and you only see that gap when all four sit side by side.

Record these as numbers you can trend month over month, not as a vibe from reading a few answers. The point of the schema below is that any two people on your team score the same answer the same way.

  • Mention rate (visibility) The share of runs for a prompt where your brand appears at all. If your brand shows up in 6 of 20 runs of "best CRM for small teams", that is a 30% mention rate, and that single number is the backbone of AI visibility.
  • Position Where you appear inside the answer: first named, in the top three, or buried at the bottom. Being mentioned tenth in a list a buyer skims is close to being invisible, so position weights the raw mention.
  • Sentiment How the engine frames you: recommended, listed neutrally, or flagged with a caveat like price or thin support. Sentiment is where reputation problems surface, because an engine can mention you often while quietly steering buyers elsewhere.
  • Share of voice Your mention rate set against named competitors on the same prompts. This is the number leadership cares about, because it says whether ChatGPT and Gemini treat you as a category leader or an also-ran.

Why do single runs lie about your AI visibility?

Single runs lie because large language models are non-deterministic: ask ChatGPT or Gemini the same question twice and you can get two different brand lists. The models sample from a probability distribution over possible next words, so the same prompt produces different answers across runs, sessions, and dates. One flattering answer is not evidence you are winning, and one bad answer is not evidence you are losing.

Treat each prompt like a poll, not a fact lookup. If you asked five people on the street to name a good running shoe, you would not publish the first reply as the market view. The same discipline applies here. A brand that appears in 2 of 20 runs and a brand that appears in 18 of 20 look identical if you only ever run each prompt once.

As a working rule, run each prompt at least 10 to 20 times per engine per cycle, in fresh sessions with memory and personalisation off. Fewer runs and your mention rate swings on noise. More runs tightens the estimate but costs time, which is exactly the trade-off that pushes teams toward automation later. Also vary nothing else: same wording, same engine settings, so the only thing moving is the model's own sampling.

Gartner expects traditional search volume to fall 25% by 2026, which is why a trustworthy AI visibility number is becoming a board-level metric.(Gartner, 2024)

How do you build a prompt set that reflects real buyers?

Build your prompt set from questions real buyers actually ask, mined from Reddit, Quora, marketplace reviews and support tickets, rather than the polished phrases your marketing team would write. Buyers do not search "leading omnichannel commerce platform". They ask "cheapest way to sell on Instagram and my own site" or "is X worth it for a small kitchen". Track the second kind, because that is what engines are answering.

Anchor the set to buying intent, not brand vanity. Prompts that name your brand only tell you how engines describe you when asked. The prompts that decide revenue are category prompts where your name has to earn its place: "best", "cheapest", "alternative to a competitor", "good for a specific use case". Keep the wording fixed once chosen, because changing a prompt resets its trend line.

Good prompt sets also model your content well. Structured, well-cited pages earn up to 40% more AI citations (Princeton GEO study, KDD 2024), so the same buyer questions you track are the ones your content should answer clearly on-site.

  • Mine, do not guess Pull real phrasing from Reddit threads, Quora questions, and Amazon or app-store reviews in your category. The awkward, specific wording is the signal.
  • Cover the buying journey Include category discovery ("best tools for X"), comparison ("A vs B"), and objection prompts ("is X worth the price"). Each stage exposes a different weakness.
  • Segment by use case and geography "Best CRM" and "best CRM for a 3-person agency in India" return different brands. Track the segments where you actually compete.
  • Freeze the wording Lock each prompt once it is in the set. If you must reword one, treat it as a new prompt with a fresh baseline, so your trend stays honest.

What is a simple monthly method a team can follow?

A team can track AI visibility with a spreadsheet and a fixed monthly routine. The method below needs no engineering. It needs discipline: same prompts, same run counts, same scoring rules, on the same week each month. Consistency is what makes month two comparable to month one.

  • Step 1: Fix your prompt set Choose 15 to 30 buyer prompts mined from real chatter. List them in a sheet with a stable ID per prompt so results line up across months.
  • Step 2: Choose your engines Pick the engines your buyers use, for example ChatGPT, Gemini, Perplexity and Google AI Overviews. Test in fresh sessions with personalisation and chat memory switched off.
  • Step 3: Run each prompt 10 to 20 times per engine Paste each prompt, record the answer, repeat. Yes, it is repetitive. That repetition is the measurement, because it converts non-determinism into a rate you can trust.
  • Step 4: Score every run For each run log four columns: mentioned yes or no, position (1st, top 3, or lower), sentiment (positive, neutral, negative), and which competitors appeared. Agree the rules up front so two people score identically.
  • Step 5: Roll up the metrics Per prompt and per engine, compute mention rate, average position, sentiment split, and share of voice versus named competitors. These four roll-ups are your scorecard.
  • Step 6: Repeat on the same week monthly Re-run the identical set next month and compare. Rising mention rate and improving position mean your AEO and GEO work is landing. Flat or falling means revisit content and sources.

When should you use a tool instead of a spreadsheet?

Switch to a tool when the manual maths stops being feasible: many prompts, times several engines, times 10 to 20 runs, times monthly, quickly runs into thousands of answers to collect and score by hand. At that scale a spreadsheet breaks on human time and consistency, and scoring drift creeps in as different people read answers differently.

Platforms exist to automate exactly this loop. Tools like Otterly and Peec track prompts across engines and chart mention rate and share of voice over time. Depra AI runs the same fixed-prompt, multi-run method across seven engines, adds SKU-level product tracking for D2C catalogues, and mines prompts from real buyer chatter rather than guesses. Each is a fair choice depending on whether you want self-serve software or a service that ships the fixes too.

The honest version: start manual for one quarter if you have never measured this. It teaches you what the metrics mean and which prompts matter. Once you know the routine is worth keeping, a tool removes the tedium and the human error, and frees your team to act on the report instead of assembling it.

Structured, well-cited pages earn up to 40% more AI citations, so tracking is only half the job: the other half is fixing the pages engines cite.(Princeton GEO study, KDD 2024)

Frequently asked

How many times should I run each prompt?

Run each prompt at least 10 to 20 times per engine per cycle, in fresh sessions with memory and personalisation off. Because AI answers are non-deterministic, fewer runs let noise masquerade as a trend. More runs tighten the estimate at the cost of time, which is the trade-off that eventually justifies automation.

Why do I get different answers to the same prompt?

Large language models sample from a probability distribution over possible next words, so ChatGPT, Gemini and others can return different brand lists for the identical question across runs and dates. This is expected behaviour, not a glitch. It is exactly why you measure a mention rate across many runs rather than trusting a single answer.

What is a good mention rate?

There is no universal threshold, because it depends on your category and the specific prompt. Judge it two ways: against your own trend month over month, and against your named competitors on the same prompts. A 30% mention rate is strong if rivals sit at 10% and weak if they sit at 70%.

Which engines should I track?

Track the engines your buyers actually use. For most consumer brands that means ChatGPT, Google AI Overviews, Gemini and Perplexity, with Claude, Copilot and DeepSeek added for fuller coverage. Test each in fresh sessions with chat memory and personalisation switched off so results stay comparable across months.

Can I track this in a spreadsheet?

Yes. A spreadsheet with one row per run and columns for mentioned, position, sentiment and competitors present handles a small prompt set well. It stops scaling when prompts times engines times run counts reach thousands of answers a month, at which point a tool removes the manual maths and the scoring drift.

How is this different from SEO rank tracking?

SEO rank tracking assumes one stable answer per query, so a single check is reliable. AI answers vary run to run, so a single check is misleading. The unit of measurement shifts from a fixed position to a mention rate across many runs, and you also score sentiment and share of voice, which classic rank tracking ignores.

How often should I re-run the tracking?

Monthly is a sensible default for most brands, run on the same week each month so cycles are comparable. Re-run more often around a big content push, a launch, or a competitor move, since those are the moments an AI visibility shift is most likely and most worth catching early.

Does Depra AI do this measurement for me?

Yes. Depra AI runs a fixed set of buyer prompts across seven engines with enough runs to be trustworthy, then reports mention rate, position, sentiment and share of voice against competitors, plus SKU-level tracking for D2C catalogues. The starting audit is free, human-delivered, and returns in five days with no card required.

Get your AI visibility measured, free

We run your category's buyer prompts across seven AI engines, with enough runs to be trustworthy, and send back a ranked report in five days. No card, no lock-in.

Keep reading