AI share of voice: how to measure it honestly (and what to do with it)

AI share of voice has become the headline metric for AI search. It’s also one of the easiest numbers in marketing to get wrong, often in ways that flatter whoever’s reporting it. This playbook gives you a version you can defend in front of a CFO.

What it is

Classic share of voice compares your visibility with competitors’ across a market: ad impressions, search rankings, media mentions. AI share of voice applies the same idea to AI-generated answers: across the questions your buyers ask, how often do answers include you rather than a competitor?

Why most AI share-of-voice numbers mislead

The hidden denominator

“We have 34% share of voice” means nothing without knowing 34% of what. The figure depends entirely on which prompts were chosen, how many, which engines, which regions and how many times each was run. Pick friendlier prompts and the number rises with no change in reality.

Rule: never report the number without the prompt set and run count beside it.

Answers aren’t stable

In SparkToro’s study with Gumshoe, 600 volunteers ran 12 prompts through three AI tools a combined 2,961 times. There was less than a 1 in 100 chance of getting the same list of brands in any two responses, and roughly a 1 in 1,000 chance of the same order. (One author has since joined an AI-tracking vendor, and the study isn’t peer reviewed, but the direction matches what anyone running repeated prompts sees.)

That has two consequences:

  • A single run tells you almost nothing.
  • Position inside an answer is close to meaningless. Appearance rate across many runs is what you can measure.

A method you can defend

1. Build the prompt set deliberately

  • Start from real buyer questions: sales calls, support tickets, search console queries, community threads.
  • Cover the funnel: problem questions (“how do I…”), category questions (“best tools for…”) and comparison questions (“X vs Y”).
  • Write prompts the way people actually ask, not stuffed with your brand’s terms.
  • Fix the set for a quarter. Add new prompts as a separate cohort so trends stay comparable.

A practical size is 30–100 prompts per market and language.

2. Run repeatedly and record everything

For each prompt and engine, run multiple times per measurement period and store the full answer text, the cited links, the engine and model version shown, the date and the region.

3. Measure four things, not one

Measure Question it answers
Mention rate How often is the brand named at all?
Citation rate How often is one of your pages linked as a source?
Recommendation rate How often is the brand suggested as an option to choose?
Narrative accuracy When you’re described, is it accurate and current?

These move independently. A brand can be mentioned often but rarely cited (engines know you, but don’t trust your pages), or cited often but described inaccurately (your pages are used, but your claims are out of date).

4. Report rates with honesty about uncertainty

Report each measure as a rate per prompt group, compared with named competitors on the same set. With small samples, say so, and treat small movements as noise until they persist. When an engine changes its model or search behaviour, annotate the chart; a jump on that date may have nothing to do with your work.

Turning the numbers into work

A metric is only useful if it changes what you do. Map each gap to a type of work:

What you see Likely cause The work
Low mention rate Weak presence in the sources engines read PR, expert content, reviews, community presence
Mentioned but rarely cited Your pages don’t answer the question well Direct, sourced, current answer pages
Cited but described inaccurately Outdated claims on your site or others’ Content audit and corrections
Strong on problem prompts, weak on comparisons No honest comparison content Comparison pages that say who each option suits
Competitor recommended for your core use case Their evidence is more specific Original data, named frameworks, case evidence

Each row is a content mission: a researched angle carried into the formats that will reach the sources engines rely on.

Connecting it to the business

AI visibility matters because of what follows it. Pew found that only about 1% of visits to pages with an AI summary involved a click on a link inside the summary, so being cited isn’t the same as being visited. Pair AI share of voice with downstream measures: branded search, direct traffic, qualified enquiries and sales conversations where buyers say they “asked ChatGPT”. The goal is influence on decisions, not a dashboard number.

Quarterly rhythm

  1. Week 1: Refresh the prompt set and run the baseline.
  2. Weeks 2–10: Work the gaps, in priority order, as missions with owners.
  3. Week 11: Re-run the same prompt set with the same method.
  4. Week 12: Report change by measure and prompt group, noting engine changes, and decide next quarter’s priorities.

For how to win the citations themselves, see how to get your brand cited by ChatGPT and other AI answers.

Questions people ask

How do you calculate AI share of voice?

Divide the number of AI answers (across a fixed prompt set and repeated runs) that mention your brand by the total number of answers, then compare that rate with competitors measured the same way. Always publish the prompt set and number of runs alongside the figure.

What is a good AI share of voice?

There's no universal benchmark, because the number depends entirely on the prompt set. Judge it against competitors on the same prompts, and against your own trend over time.

Why does my AI share of voice change week to week?

Partly because AI answers vary between runs and partly because engines update their models and search behaviour. Small changes are often noise; only trust movements that persist across repeated runs.

Is position in an AI answer meaningful?

Rarely. In SparkToro's study, two answers to the same prompt listed brands in the same order only about once in 1,000 runs, so appearance rate is a far more reliable measure than position.

Sources

  1. AIs are highly inconsistent when recommending brands or products, SparkToro
  2. Google users are less likely to click on links when an AI summary appears, Pew Research Center

Facts on this page were last checked on . First published 11 October 2026.

Arun Bansal

Founder, ReachFabric

Arun Bansal has built and run infrastructure and software companies since 2009: ServerGuy (merged into ZenoCloud in 2024), ZenoCloud, which manages cloud, security and AI infrastructure for 170+ businesses, and Breeze.io, an AI-first website builder. Writes Zeno Briefs every week.

LinkedIn

Want this run for you, with evidence attached?

Bring one content problem from this playbook. We will scope a pilot that runs it end to end.

Starts at $2,999/month