Guides / AI share of voice
How to measure AI share of voice against competitors in ChatGPT and Gemini: the B2B method
By Olly, founder of AEOscore, AEO consultant and full-stack developer, 20+ years in search16 min read
AI share of voice is the percentage of all brand mentions in AI answers that belong to your brand rather than a competitor, measured across a fixed set of buyer questions. It is a competitive, zero-sum number: your share can only rise if somebody else's falls, which is what makes it worth reporting.
AI share of voice = (your brand's mentions ÷ total brand mentions in the same tracked answers) × 100
TL;DR
- Count every brand named across a frozen set of 30 to 100 buyer prompts, run at least three times per engine; your mentions divided by all brand mentions, times 100, is your share of voice.
- Report ChatGPT and Gemini separately, count every brand the engines name (not just your chosen rivals), and treat any single run as a sample, never as the answer.
Read the video transcript
AI share of voice. How much of the AI answer belongs to you?
The formula: share of voice equals your mentions divided by all brand mentions, times 100. Count every brand named. Divide. That is the number.
Worked example: 100 prompts, three runs each, 480 brand mentions. Your brand: 62 mentions, 12.9%. Competitor A: 118 mentions, 24.6%. 62 of 480 mentions is 12.9%, and the gap to close is 11.7 points.
Count every brand the engine names, not just the rivals you chose. The denominator you leave out is exactly where the flattering lie lives.
Three of the seven traps: dividing by prompts instead of mentions, a prompt set that keeps changing, one run treated as the answer. The guide lists all seven.
Freeze the prompt set. Average the runs. We asked one buyer question 40 times; no two answers matched (our own variance test, 20 August 2026).
Measuring is the easy part. The job is moving it: publish what the engines could not find, get named where they read, re-scan.
Read the B2B method: aeoscore.co.uk/guides/ai-share-of-voice.
On this page
Definitions
What AI share of voice is (and what it is not)
What is AI share of voice?
AI share of voice is the percentage of all brand mentions in AI answers, across a fixed set of buyer questions, that belong to your brand rather than a competitor.
Is share of voice the same as AI visibility score?
No: share of voice is your proportion of total brand mentions, while visibility score is the percentage of prompts in which you appear at all, so a brand can have a high visibility score and a low share of voice.
Three numbers get run together in this space, and mixing them up sends six months of work at the wrong target. Share of voice, visibility score and citation share measure different things and answer different questions:
| Metric | What it counts | The question it answers | What to use it for |
|---|---|---|---|
| Share of voice | Your share of every brand mention across the answer set | Of all the brand talk, how much is about us? | The competitive number for a board or client: zero-sum, comparable month to month |
| AI visibility score | The percentage of prompts where you appear at all | Did we turn up? | The coverage number: finds the questions where you are absent entirely |
| Citation share | How often your domain is linked as a source | Is our site doing the evidencing? | The sourcing number: shows whose pages the engines actually read |
A brand can hold a strong visibility score and a weak share of voice at the same time: you appear in most answers, but always mid-list behind two rivals who soak up the mentions. The reverse happens too. Track both; report them side by side. The deeper definitions live in the glossary under share of voice, AI visibility score and prompt set.
“Share of voice is your slice of the mentions; visibility score is whether you turned up at all. Confuse the two and you will optimise the wrong number for months.”
Why the measurement matters now, with dates:
- 94% of business buyers report using AI in their buying process (Forrester Buyers' Journey Survey 2025, published January 2026) Forrester
- ChatGPT passed 900 million weekly active users in February 2026; the Gemini app passed 1 billion monthly users in August 2026 (different metrics, both vendor-announced) OpenAI
- Only 6.8% of ChatGPT answers included an outbound link as of May 2026 (Similarweb, US desktop), which is why counting linked citations alone undercounts you badly there Search Engine Journal
- Across 2,961 runs of 12 identical prompts, the chance of two runs returning the same brand list was under 1 in 100 (SparkToro, January 2026) SparkToro
The maths
The formula, with a worked B2B example
How do I calculate share of voice in ChatGPT and Gemini?
Divide the number of times your brand is named across your tracked answer set by the total number of brand mentions in that same set, including every competitor named, then multiply by 100.
Worked through with real numbers: a B2B brand tracks 100 buyer prompts, runs each 3 times per engine, and counts every brand named across the answers, its own 6 tracked rivals plus anything else the engines bring up. The answer set contains 480 brand mentions in total:
| Brand | Mentions | Share of voice |
|---|---|---|
| Your brand | 62 | 12.9% |
| Competitor A | 118 | 24.6% |
| Competitor B | 96 | 20.0% |
| Competitor C | 74 | 15.4% |
| Competitor D | 58 | 12.1% |
| Competitor E | 47 | 9.8% |
| Competitor F | 25 | 5.2% |
| Total | 480 | 100% |
The denominator is mentions, not prompts. Dividing 62 by the 100 prompts would produce 62%, a number that describes nothing. The division is always by the 480 brand mentions in the same answer set.
Inputs
How to build your prompt set
How many prompts do I need to track?
Between 30 and 100 buyer-intent prompts is enough for a stable B2B figure, and the set must stay frozen so your trend line means something.
Build the set from questions buyers actually ask, not from the keywords you rank for. Four types cover the B2B journey; a balanced set draws on all of them:
Problem-led
The buyer describes the pain, not the product. Engines answer with categories and named providers, so these prompts catch you at the earliest moment a shortlist forms.
- “Our website enquiries have dropped but our Google rankings look the same. What is going on?”
- “Why does ChatGPT recommend our competitors but not us?”
Category discovery
The buyer knows the category and asks who is in it. These are the prompts where share of voice is won and lost, because every answer is a list of brands.
- “What is the best CRM for a UK manufacturing SME?”
- “Which payroll providers should a 50-person UK company consider?”
Comparison
The buyer has a shortlist and wants it whittled. Engines lean on comparison pages and reviews here, so these prompts test whether the comparing content mentions you.
- “Xero vs QuickBooks for a UK limited company: which is better?”
- “What are the best alternatives to Salesforce for a small B2B team?”
Direct recommendation
The buyer asks the engine to choose. The answer usually names one to three brands with reasons, which makes these the highest-stakes prompts in the set.
- “Recommend an IT support provider for a 30-person firm in Manchester.”
- “I need a B2B telemarketing agency in the UK. Who should I use?”
Then freeze the set. Every added, removed or reworded prompt changes the denominator, and a changed denominator means this month’s figure no longer compares with last month’s. If the set must change, restate the baseline and start the trend line again from that date. Thirty prompts spread across the four types, written in your buyers’ own words, is a solid starting set.
Inputs
Choosing the competitor set
Which competitors should I include in the count?
Include the 3 to 8 rivals your buyers genuinely shortlist, but count every other brand the model names as well, because leaving brands out of the denominator inflates your own score.
Name 3 to 8 rivals a buyer would genuinely shortlist against you, and track those by name. Then apply the open-pool rule: every other brand an answer names goes into the denominator too. The engines routinely recommend businesses you have never considered rivals, and occasionally skip the ones you watch most closely; both facts belong in your number. In our own scans, answers to category questions regularly name brands the client had never listed. Across the latest weekly scans of the three UK B2B sites we monitor (490 scored answers, August 2026), the median answer names 2 distinct brands, and category-discovery prompts routinely name eight or more.
“Count every brand the engine names, not just the rivals you chose. The denominator you leave out is exactly where the flattering lie lives.”
Per engine
Why ChatGPT and Gemini give different answers
Should I measure ChatGPT and Gemini separately or together?
Measure them separately and report them separately, because ChatGPT frequently answers without citing any source while Gemini leans on retrieved web pages, so a blended average hides which engine you are actually losing.
The two engines build answers differently, and the difference decides how you must count. Similarweb’s 2026 Generative AI Landscape report measured that only 6.8% of ChatGPT answers included an outbound link as of May 2026, up from 1.3% a year earlier (US desktop, reported by Search Engine Journal). Semrush’s clickstream study found ChatGPT switched on web search for 34.5% of queries in February 2026, with the rest answered from the model’s own memory (Semrush). Measure ChatGPT by linked citations alone and most of your presence there is invisible to you.
Gemini sits at the other end. Google’s own documentation describes Gemini answers grounded in live Google Search results with sources attached (Google), and Google AI Overviews and AI Mode are built from the Search index by construction. Mentions there track the retrievable web much more closely, which is why a fix that moves Gemini often does nothing for ChatGPT, and the other way round.
Even when ChatGPT does retrieve a page, Ahrefs’ study of 1.4 million prompts found it cites that page in the answer only about half the time (Ahrefs). Getting read is not the same as getting credited.
So report each engine on its own line, and keep any blended figure as a secondary number. A blended 14% that hides 22% on Gemini and 6% on ChatGPT does not tell you where the problem lives; the per-engine lines do. The same goes for the other four engines a full measurement covers: our methodology page sets out how we collect ChatGPT, Claude, Gemini, Perplexity, Google AI Overviews and AI Mode without averaging them into mush.
Method
Running the scans properly
How often should I run the scans?
Weekly is the minimum for a reliable trend, with daily runs on your highest-value prompts, because AI answers vary between runs even when nothing has changed.
The mechanics matter as much as the cadence. Fresh session per prompt. Memory and custom instructions off, logged out where the engine allows it, because a logged-in account that has read your chats is the least representative sample available to you. Locale set to the UK, since UK-worded prompts surface different brands than US ones. The same day each week. And several re-asks per prompt, averaged, because a single run is a sample of a distribution, not the distribution.
Can I measure AI share of voice for free in a spreadsheet?
You can log answers, brands and dates by hand against a frozen prompt set, but three runs per prompt across even two engines means logging around 180 answers a week, and most teams find that unsustainable within a month.
The variance is not a rounding error. SparkToro ran 12 identical prompts 2,961 times across ChatGPT, Claude and Google’s AI in late 2025 and found under a 1 in 100 chance that two runs return the same brand list; the same study found aggregate frequency across many runs far more stable, which is the whole case for measuring this way (SparkToro).
We ran the test ourselves
On 20 August 2026 we asked one buyer prompt, “What are the best AI visibility monitoring tools for a UK B2B company?”, 20 times each on ChatGPT and Gemini, every run in a fresh session, and extracted every brand named. ChatGPT returned 20 different brand line-ups across 20 runs, naming 113 different brands in total, up to 23 in a single answer. Gemini returned 20 different line-ups, naming 58 different brands, up to 17 in one answer. No two of the 40 runs produced the same brand list. A single manual check in your own account is one draw from that pool, and it will mislead you in whichever direction it happens to land.
Pitfalls
The seven traps that ruin the number
1. Closed-pool denominator
Counting only the rivals you chose to track inflates your share: the engines name brands you have never heard of, and those mentions belong in the denominator. Count every brand named.
2. Dividing by prompt count instead of mention count
62 mentions across 100 prompts is not 62% of anything. Share of voice divides by total brand mentions (480 in our example), not by the number of prompts (100); mixing the two produces a flattering nonsense number.
3. Prompt-set drift
Add, remove or reword prompts and the denominator changes, so this month's figure no longer compares with last month's. Freeze the set; if you must change it, restate the baseline and start the trend line again.
4. Model version changes
Engines swap underlying models without telling you, and the answers change with them. Log the date of every run so a sudden shift can be read against a model change rather than against your own work.
5. Personalisation
Your logged-in account, with memory and custom instructions on, is the least representative sample available to you. It has read your chats; your buyers' accounts have not.
6. Single-run sampling
The same prompt returns different brand lists on different runs even when nothing has changed. Three or more runs averaged gives you a usable figure; a single run tells you very little.
7. Mixing engines into one blended number
ChatGPT and Gemini disagree often enough that an average hides which one you are losing. Report per engine first; keep any blended figure as a secondary line.
From number to fixes
What to do when the number is low
How do I improve share of voice once I have measured it?
Fix the gaps individually: publish the comparison, pricing and specification detail the engines could not find, get named in the third-party roundups and listicles the engines are quoting, and re-scan to confirm the answer changed.
A share-of-voice number on its own changes nothing, so here is each kind of gap mapped to its fix:
Not mentioned anywhere
The engines have nothing to say about you. Publish the pages that give them sentences to lift: plain answers to the buyer questions, a public pricing page, comparison pages that name rivals fairly. Then get into the third-party roundups and directories the engines cite, because absence there is usually the root cause.
Mentioned but not recommended
You appear mid-list without a reason attached. Give the engines a differentiator they can repeat: a specific claim with a number or date on your own pages, and reviews or case write-ups that say the same thing in a third-party voice.
A competitor is named instead of you
Read the citations on those answers. The fix is usually to be present on the exact pages doing the recommending: the roundup that lists them, the comparison page they wrote, the thread where they were praised. Pitch the roundup, publish the comparison, join the thread honestly.
Cited or described inaccurately
Wrong prices, retired products, a mixed-up namesake. Correct the source pages the engines read, state the facts plainly on your own site (a facts page works), and re-scan until the answer changes; inaccuracy is the one gap where the fix is fully in your hands.
“Measuring share of voice gets you the number. Moving it is the job: publish what the engines could not find, get named where they read, and re-scan.”
The one-pager
Reporting it to a board or client
What is a good AI share of voice for a B2B brand?
There is no universal benchmark, so set your baseline from your first month of runs and target beating the highest-scoring competitor in your own tracked set.
The monthly one-pager that keeps a board interested has six rows and no dashboard screenshots:
- One line per engine: share of voice this month, last month, and the movement, with your top competitor's figure beside it
- The blended figure once, beneath the per-engine lines, marked as secondary
- Visibility score alongside, so a rising share on falling coverage cannot hide
- The three biggest answer-level changes, quoted: what the engine said, which prompt, which engine
- The fix log: what shipped since last month, and which prompt each fix targets
- Next month's three fixes, each tied to a named gap
The pairing of movement with the fix log is the part that earns the budget: it shows the number moving because of named work, not weather. Share of voice is one of nine KPIs in the full reporting framework; the rest, with a downloadable monthly template, are in the agency guide to AEO KPIs and client reporting.
Who we are
How AEOscore measures it
AEOscore runs this exact method as software with a person attached: a frozen panel of UK-phrased buyer prompts (starting at 25, ceiling 40 active), scanned weekly (daily on the top tier) across ChatGPT, Claude, Gemini, Perplexity, Google AI Overviews and AI Mode, every answer scored for mentions, prominence and sentiment, share of voice and visibility score reported per engine, and a written fix plan for every gap the scan finds. On the top tier the fixes are made for you by a named in-house strategist and developer. Built and run in the UK by Goreblimey Ltd, Cheltenham.
Pricing in one line: £395 one-off AI Visibility Audit, credited against your first month, £99 a month Monitor, £1,495 a month Monitor + Strategist. Prices exclude VAT. Details on the pricing page, the scoring maths on the methodology page, and the free AEO score will tell you your starting point across all six engines before you spend anything.
FAQ
AI share of voice: common questions
- Do brand mentions without a link still count?
- Yes: a mention with no citation is still a recommendation a buyer reads, which is why counting only linked citations undercounts you badly on ChatGPT in particular.
- Can I use the free ChatGPT plan for the measurement?
- You can, but log out or use a temporary chat so memory and custom instructions cannot personalise the answers you are trying to measure.
- What causes a sudden drop in share of voice when we changed nothing?
- Usually a model or retrieval change on the engine's side, or a rival landing in a heavily-cited roundup; check your run log against the date and read the citations on the changed answers before reacting.
- Should branded prompts be in the share of voice set?
- Keep them in the wider tracked panel but out of the share of voice denominator, because questions containing your own name inflate your share without telling you anything competitive.
- Is asking through the API the same as asking in the app?
- Not always: API answers can differ from what a consumer sees in the app, which is one reason to treat every figure as directional and to keep your method consistent rather than mixing collection routes.
- Does AI share of voice predict revenue?
- Not directly: it is a leading indicator of shortlist presence, so pair it with branded search volume and enquiry sources rather than reporting it as pipeline.
Sources
Numbered sources
- Forrester (John Buten), B2B Buyers Make Zero-Click Buying Number One. Accessed 20 August 2026. 94% of business buyers report using AI in their buying process (Forrester Buyers' Journey Survey, 2025). Published 22 January 2026.
- OpenAI, Scaling AI for everyone. Accessed 20 August 2026. More than 900 million weekly active ChatGPT users. Published 27 February 2026.
- Google, The Gemini app surpasses 1 billion monthly users. Accessed 20 August 2026. 1 billion monthly active users for the Gemini app. Published 11 August 2026. Weekly and monthly actives are different metrics; the two figures are not directly comparable.
- Search Engine Journal, reporting Similarweb's 2026 Generative AI Landscape report, ChatGPT Links Out Most On Travel Queries, Data Shows. Accessed 20 August 2026. 6.8% of ChatGPT answers included outbound links as of May 2026 (up from 1.3% in June 2025). US desktop usage; Similarweb describes its figures as estimates. Published 27 July 2026.
- Semrush, ChatGPT traffic analysis: insights from 17 months of clickstream data. Accessed 20 August 2026. ChatGPT enabled web search on 34.5% of queries as of February 2026, with month-to-month swings between 15% and 66.3%. US clickstream panel. Published 7 April 2026.
- SparkToro (Rand Fishkin), AIs are highly inconsistent when recommending brands or products. Accessed 20 August 2026. 2,961 runs of 12 identical prompts: under a 1 in 100 chance that two runs return the same brand list, while aggregate frequency across many runs is far more stable. Published 28 January 2026.
- Ahrefs (Louise Linehan), Why ChatGPT Cites One Page Over Another (Study of 1.4M Prompts). Accessed 20 August 2026. Of the URLs ChatGPT retrieves, 49.98% get cited in the answer. A per-retrieved-URL figure, not a per-answer one. Published 15 April 2026.
- Google, Grounding with Google Search (Gemini API documentation). Accessed August 2026. Google's own description of Gemini answers grounded in live Google Search results with source attribution.
- AEOscore, Run-to-run variance test and fleet figures. Accessed 20 August 2026. Method described on this page; measured 20 August 2026 from live engine runs and the AEOscore scan database.
What changed on this page
- : Guide published: formula, worked example, prompt-set and competitor-set rules, per-engine differences with dated sources, the seven traps, and our own 40-run variance test.
Know your starting point first
The free AEO score measures you across all six engines, human-checked, within 1 to 2 working days. Domain and email, no card.
- ChatGPT
- Claude
- Gemini
- Perplexity
- Google AI Overviews
- Google AI Mode