AI share of voice: measure it by hand with a fixed prompt set

AI share of voice is the fraction of a fixed prompt set whose answers name you. Here is how to build the list, the seven-column record sheet, and the four things that invalidate a comparison.

Measurement7 min read1250 views
AI share of voice: measure it by hand with a fixed prompt set

AI share of voice is the fraction of a fixed set of prompts whose answers name you. You can measure it by hand, in about 40 minutes a month, with a prompt list you write once and never edit. This chapter gives the rules for building that list, the record sheet to put the results in, and the four things that quietly invalidate a comparison between two rounds.

Read this first

You need to know which questions matter to you before you can count anything, and that decision comes from what people actually search, not from what you wish they asked. If you have not been through that, do it first — the shortlist you produce there becomes half of the prompt set here.

One term, defined so the rest of the chapter holds: an answer names you when your product or company appears in the text of a generated answer. That is different from being cited, where the engine attributes a sentence to a URL you own. This chapter counts names. Both get recorded, and they are separate columns.

Why AI share of voice cannot be a rank tracker

Rank tracking works because a search result page is a list with positions, produced by a system that returns nearly the same list to the same request. An answer has none of those properties. There is no position three. Ask the same thing twice and the wording changes, sometimes the recommendations change, and you have no version number to attribute the difference to.

So the measurement has to give up precision to get repeatability. Instead of "where do we rank", the question becomes "in how many of these 20 answers does our name appear at all". That is a coarse number. It is also stable enough that a real change shows up in it, which is the only property that matters for deciding whether something worked.

Measure the same 20 questions forever. The number is only worth anything as a difference against itself.

How to build the prompt set

The prompt set is the instrument. Everything else in this chapter is bookkeeping, and a bad instrument cannot be fixed later by careful bookkeeping.

RuleWhy
20 prompts, written before the first runAt n=20 one prompt moves the result by 5 points, which is coarse but readable
Fixed wording, never editedEditing a prompt resets the baseline and you lose the comparison
Phrased the way a person would askYou are testing the question you want to win, not a keyword
No brand name in 15 of themAsking "what is your product" guarantees a mention and measures nothing
One turn only, fresh sessionFollow-ups and chat memory make two rounds incomparable

For composition, spend most of the list on questions where you have no automatic advantage. A workable split is 15 prompts with no brand name — the category questions, the job-to-be-done questions, the "X versus Y" questions — and 5 that name you, which measure something different: whether the engine describes you correctly when it already knows who you mean.

Write the list in a file, commit it, and treat edits as a version bump. When you do need to change a prompt, add the new one and keep the old, rather than replacing it. The moment your denominator changes, so does every number you have collected.

Running a round

A round is one pass through the list on each engine you track. Three to four engines is plenty; more engines cost time and add nothing until the first three disagree.

  1. Open a fresh session, logged out where the engine allows it, and disable any memory or personalisation setting you can see.
  2. Paste each prompt exactly as written, take the first answer, and do not follow up.
  3. Record the row before moving on. Reading twenty answers and filling the sheet afterwards produces sheets that agree with your hopes.
  4. Save the raw answer text somewhere, even if only in a folder of files named by prompt number and date.

Step four is what turns this from a metric into evidence. Six months later, the number tells you something changed; the saved answers tell you what the engine actually said, and that is the part you can act on.

The deliverable: the record sheet

One row per prompt per engine per round. Seven columns.

ColumnWhat goes in it
DateThe date of the round, not of the analysis
EngineProduct name as the user sees it
Prompt IDThe number from your fixed list
NamedYes or no — did your name appear in the text
First namedYes or no — were you the first product mentioned
Others namedEvery other product named, comma separated
Cited URLThe URL of yours it linked, or blank

From those columns you get three figures per engine: share of voice (named ÷ prompts), first-name share (first named ÷ prompts), and a competitor tally that you did not have to guess at. The competitor column is usually the most useful one in the first six months, because it tells you which set of products the engine thinks you belong with.

How often to re-run

Monthly is the floor and roughly the right cadence. Weekly produces movement you cannot attribute to anything, because the engines change under you at their own pace and you will read that noise as a result.

The exception is a deliberate change. When something ships that is supposed to affect this — a rewritten page, a new integration listing, a burst of off-site mentions — run an extra round 14 days after it lands, and compare it against the round before it shipped rather than against the last monthly one.

Four things that invalidate a comparison

All four produce a number that looks fine and means nothing.

  1. Edited prompts. Even a word. The list is the denominator.
  2. A logged-in session. Personalisation and chat memory change answers, and you cannot subtract them afterwards.
  3. A different country. Answers vary by region; if you run one round from a laptop abroad, that round belongs to a different series.
  4. A different engine version. This one you cannot control or even see, which is why the boundary section below exists.

What this measurement cannot tell you

It cannot tell you why the number moved. Engines change their models and their retrieval layers without notice, and from outside there is no way to separate "our work landed" from "the engine changed". Two rounds in a row moving in the same direction is weak evidence. One round is not evidence at all.

It also does not scale down. At 20 prompts, a single prompt is 5 percentage points, so a move from 40% to 45% is one answer changing its mind. Treat anything under 10 points as unresolved rather than as a result.

And we have not validated this method against any vendor's tracking product. We do not know how their numbers are produced, so we make no claim that ours matches theirs. What this procedure gives you is a series you produced yourself, with a documented instrument, which is a different and more defensible thing than a number you cannot audit.

What to do with the result

A low share of voice with your competitors named repeatedly is a content and off-site problem, and the off-site half of it is how to get mentioned by AI. A share of voice that is fine while the descriptions are wrong is a different problem, and it is fixed on your own pages, starting with what makes an engine choose a source in the first place — see how AI engines pick sources.

Either way, run round one before you change anything, and file it with the rest of your starting values as described in recording a baseline before you change anything. Turning this loop into something that runs without a person is what QueryWin is being built to do.

Common questions

How do you measure AI share of voice without a tool?

Twenty fixed prompts, three or four engines, one fresh session per prompt, seven columns per row, once a month. The whole round takes about 40 minutes once the list exists.

How many prompts is enough?

Twenty is the smallest list where the arithmetic stays readable — each prompt is 5 points. Forty halves that granularity and doubles the time. Below fifteen the number jumps around too much to compare.

Should I count a mention that has no link?

Yes, in the Named column, and record the missing link by leaving the Cited URL column blank. Being named and being cited are different outcomes with different fixes, which is why they are two columns and not one.

What if the answer names us but describes us wrong?

That counts as named, and it is the more urgent problem. A wrong description repeated across engines usually traces back to how your own pages describe you, not to anything off-site.

Part of the QueryWin handbook · Level 2

AI share of voice: measure it by hand with a fixed prompt set