Guides / AEO KPIs and client reporting
AEO KPIs that replace keyword rankings: an agency guide to reporting AI visibility to clients
By Olly, founder of AEOscore, AEO consultant and full-stack developer, 20+ years in search18 min read
Keyword rankings are replaced by visibility KPIs: mention rate, citation rate, AI share of voice, answer inclusion rate, recommendation position and sentiment, plus outcome KPIs such as AI referral traffic and assisted conversions, tracked across six engines. Rankings measured your position on one list. AEO KPIs measure how often, how favourably and how accurately AI engines describe your client.
TL;DR
- Track six visibility KPIs (mention rate, citation rate, share of voice, answer inclusion by prompt bucket, recommendation position, sentiment and accuracy) and three outcome KPIs (cross-engine coverage, AI referral traffic, gap-to-fix closure rate).
- Report monthly as rates across a frozen prompt set, never as single snapshots, and put the fix log next to the numbers so the client sees work as well as weather.
Read the video transcript
For agencies. The rankings report is measuring the wrong thing.
AI answers have no position one to ten. 68% of Google searches now end without a click (clickstream data, January to April 2026).
Nine KPIs replace it, and every one has a formula: mention rate, citation rate, share of voice, answer inclusion, position in the answer, sentiment and accuracy, cross-engine coverage, AI referral traffic, and gap-to-fix closure.
The monthly report starts with an executive snapshot: mention rate on non-brand questions, share of voice against rivals, citation rate, and the gaps fixed this month. One page a client can read: the score, the rivals, the work done.
Then the trend: mention rate month by month, charted against the fixes that caused the movement.
Report the rate, never the run. We asked one buyer question 40 times; no two answers matched (our own variance test, 20 August 2026).
The fix log sits beside the numbers. Clients pay for the work.
The full guide has every formula, and the template: nine KPIs, starting ranges from live scans, and a monthly client report you can copy.
Read the agency guide: aeoscore.co.uk/aeo-kpis-client-reporting. Free monthly client report template included.
On this page
The shift
Why rankings stopped working as a client KPI
Do keyword rankings still matter at all?
Rankings still matter for commercial, bottom-of-funnel queries where people click to buy, but for informational questions a top ranking is now a prerequisite for being cited rather than a traffic driver in its own right.
The rankings report assumed a chain: position earns the click, the click earns the visit, the visit earns the lead. AI answers break the chain at the first link. There is no position one to ten in a ChatGPT answer, and a client can hold page one on every tracked keyword while the buyers who used to click it read an AI summary and never arrive. When that happens, the rankings report says nothing changed. The traffic says otherwise, and the client notices the traffic.
The evidence, with dates:
- 68% of Google searches ended without a click to the open web in early 2026 (SparkToro and Similarweb clickstream, January to April 2026) Search Engine Land
- Organic click-through on queries showing an AI Overview fell to 0.61% by September 2025, from 1.76% in June 2024: a decline of around 61% (Seer Interactive, published November 2025) Seer Interactive
- Pew found US users clicked a traditional result on 8% of visits when an AI summary appeared, against 15% without one (March 2025 data) Pew Research Center
- 94% of business buyers report using AI in their buying process (Forrester Buyers' Journey Survey 2025, published January 2026) Forrester
- ChatGPT referrals to tracked B2B websites roughly quadrupled in a year: 645,000 monthly visits in June 2025 to 2.6 million in June 2026 (Labs by Demandbase, published August 2026) Demand Gen Report
“A ranking is a position on a list nobody may be reading. What I want to know is whether the buyer is being told my client's name.”
Old KPI to new
Every SEO KPI and its AEO replacement
The fastest way to move a client from the old report to the new one is a straight translation. Each row keeps the job the old KPI was doing and names the metric that does that job now:
| SEO KPI | AEO replacement | How to measure it | What good looks like |
|---|---|---|---|
| Keyword ranking position | Mention rate | Frozen set of 30 to 100 buyer prompts, three or more runs per engine, weekly; share of answers naming the brand | Non-brand mention rate climbing month on month; single digits at the start is normal |
| Visibility index / ranking share | AI share of voice | Your mentions over all brand mentions in the same answers, every named brand counted | Closing the gap on the named competitor directly above you in the same answer set |
| Organic click-through rate | Citation rate | Share of answers linking the client's domain, from the source lists engines attach | Citation rate growing towards mention rate; half of it is a fair start |
| Featured snippet ownership | Recommendation position | Median list position across answers that name the brand in a provider list | Median moving from mid-list towards the first two names |
| Brand SERP health | Sentiment and accuracy | Every answer naming the brand scored positive, neutral or negative, factual errors logged with evidence | Negatives diagnosed and falling; zero uncorrected factual errors older than a month |
| Organic sessions | AI referral traffic and assisted conversions | GA4 AI Assistants channel plus a retroactive custom channel group, treated as a floor | A rising monthly line with conversion rate reported beside organic's |
| Tasks completed / sprint velocity | Gap-to-fix closure rate | Fix log: gaps identified, fixes shipped, re-scan outcome per fix | A steady shipped-fix cadence with before and after answers attached |
The framework
The nine KPIs that replace keyword rankings
What KPIs replace keyword rankings in AEO reporting?
Keyword rankings are replaced by six visibility KPIs (mention rate, citation rate, AI share of voice, answer inclusion rate, recommendation position, and sentiment and accuracy) plus three outcome KPIs (cross-engine coverage, AI referral traffic and assisted conversions, and gap-to-fix closure rate).
What is the difference between a mention and a citation?
A mention is when an AI engine names your client's brand in its answer text, while a citation is when it links to or sources their domain, and a healthy account grows both, because mentions build preference and citations build referral traffic.
Each KPI below comes with its definition, the formula, how to sample it and what a realistic starting number looks like. The starting ranges are drawn from our own scans of live UK B2B sites, not invented benchmarks.
KPI 1 · Visibility
Mention rate
The share of tracked prompts where the AI engine names your client's brand in its answer.
Mention rate = answers naming the brand ÷ total scored answers in the tracked set × 100
How to sample it: Run a frozen set of 30 to 100 buyer prompts at least three times per prompt per engine, weekly, and count a prompt as a mention if any run names the brand. Report the rate per engine, per month.
Realistic starting range: On non-brand buyer prompts, expect single digits at the start: across the three UK B2B sites we scan, the median non-brand mention rate is 6%, and the spread runs from under 1% to 23%. Brand-name prompts sit near 100% and belong on a separate line.
KPI 2 · Visibility
Citation rate
The share of tracked answers that link to or source your client's domain, not just name the brand.
Citation rate = answers citing the client's domain ÷ total scored answers × 100
How to sample it: Capture the source list attached to each answer (Gemini, Perplexity and Google AI Overviews expose one; ChatGPT often does not) and match hosts against the client's domain. Same frozen prompt set as mention rate, so the two lines compare.
Realistic starting range: Always lower than mention rate, and much lower on ChatGPT: in our August 2026 scans, 46% of ChatGPT answers that named a monitored brand carried no link to its domain, against 8% on Gemini. A citation rate of half your mention rate is a normal starting point.
KPI 4 · Visibility
Answer inclusion rate by prompt bucket
Mention rate split by prompt type: brand, soft-brand and non-brand buyer questions reported as separate lines.
Inclusion rate per bucket = answers naming the brand in that bucket ÷ scored answers in that bucket × 100
How to sample it: Tag every prompt in the set as brand (contains the client's name), soft-brand (their category plus a distinctive detail) or non-brand (pure buyer question), and never blend the buckets. The blended figure flatters, because brand prompts are easy.
Realistic starting range: Brand prompts near 100%, soft-brand somewhere in the middle, non-brand in single digits at the start. The non-brand line is the one that predicts new business, so it leads the report even when it is the ugliest number on the page.
KPI 5 · Visibility
Recommendation position
Where in the answer the client appears when the engine lists providers: first named, third, or last.
Recommendation position = median list position of the brand across answers that name it in a list
How to sample it: For every answer that returns a list of providers, record the client's position in it. Track the median rather than the mean, because one stray answer that names them tenth should not swamp a month of second places.
Realistic starting range: Being named at all comes first; position work comes second. A brand new to the answers typically enters mid-list. First-named matters because assistants often summarise their own list, and the summary keeps the first two or three names.
KPI 6 · Visibility
Sentiment and accuracy
How favourably each answer describes the client, and whether the facts in the description are right.
Sentiment split = positive, neutral and negative answers as a share of answers naming the brand; accuracy = factual errors found per month, listed
How to sample it: Read (or machine-score, then spot-check) every answer that names the client, and log the quoted sentence behind each negative call so the client can see the evidence. Log every wrong price, retired product and mixed-up namesake as a named error, with the engine and prompt.
Realistic starting range: Most B2B answers are neutral; a handful of negatives is normal and fixable. Accuracy errors are common where pricing changed or a namesake exists. Every accuracy error is a work item, which is exactly why it belongs in a KPI report.
KPI 7 · Outcome
Cross-engine coverage
How many of the six engines (ChatGPT, Claude, Gemini, Perplexity, Google AI Overviews, AI Mode) mention the client on non-brand prompts.
Cross-engine coverage = engines with a non-zero non-brand mention rate ÷ engines tracked × 100
How to sample it: Same prompt set, run per engine, scored per engine. Never average engines into one blended figure: the engines source answers differently, so a fix that moves Gemini often does nothing for ChatGPT, and the per-engine lines are what tell you where the problem lives.
Realistic starting range: Most clients start visible on one or two engines and absent on the rest, because Gemini and AI Overviews lean on the live Google index while ChatGPT leans on training data plus occasional search. Six out of six on non-brand prompts is rare and worth reporting loudly.
KPI 8 · Outcome
AI referral traffic and assisted conversions
Sessions, key events and revenue from visitors arriving off AI assistant domains, measured in the client's analytics.
AI referral share = sessions from AI assistant sources ÷ all sessions × 100, with conversions counted on the same segment
How to sample it: GA4 shipped a default AI Assistants channel in May 2026: start there, then build a custom channel group on a referrer regex for chatgpt.com, perplexity.ai, claude.ai, gemini.google.com and copilot.microsoft.com, because custom groups apply retroactively and give you history. Treat the number as a floor: assistants strip referrers, and some readers type the brand into Google instead.
Realistic starting range: Small but growing fast: ChatGPT referrals to tracked B2B sites roughly quadrupled in the year to June 2026 (Labs by Demandbase). Expect a low single-digit share of sessions that converts respectably, and chart it monthly next to organic.
KPI 9 · Outcome
Gap-to-fix closure rate
The share of identified visibility gaps that were actually fixed this month, on the client's site or off it.
Closure rate = gaps fixed and shipped this month ÷ gaps identified and accepted × 100
How to sample it: Keep a fix log: every gap the scans surface, who owns it, the date it shipped, and the re-scan result afterwards. Count a fix as closed when the change is live, and separately track whether the answers moved at the next scan.
Realistic starting range: This is the KPI that separates a reporting agency from a fixing one, and the client's clearest evidence of work done. A steady cadence of a handful of shipped fixes a month, each tied to a named gap, beats a long list of identified-but-untouched items every time.
Reading them
Leading indicators vs outcome KPIs
The client’s revenue KPIs have not changed: traffic, leads, pipeline and revenue are still the scoreboard. What changed is the set of leading indicators that predict them.
Rankings were never the point; they were the leading indicator everyone agreed to watch because they predicted clicks, and clicks predicted revenue. That prediction is what broke. The six visibility KPIs are the new leading indicators: they measure whether the client is present, favourably described and evidenced at the moment a buyer asks an assistant who to shortlist. The three outcome KPIs (cross-engine coverage, AI referral traffic and assisted conversions, and the closure rate on fixes) connect that presence back to the scoreboard the client already trusts.
Say this plainly in the report and the conversation gets easier: the dial that used to predict the client’s revenue has been replaced, and this report is where they watch the new one.
Method
How to sample AI answers so the numbers are defensible
How many prompts should you track per client?
Track 30 to 100 real buyer questions per client, weighted toward prompts that sit close to a purchase decision, rather than a large volume of top-funnel prompts that fluctuate and do not correlate with revenue.
Build the set from questions the client’s buyers actually ask: problem descriptions, category discovery, comparisons and direct recommendation requests, written in the buyer’s words, with UK phrasing if the client sells in the UK. Tag each prompt brand, soft-brand or non-brand. Then freeze the set: every added or reworded prompt changes the denominator, and a changed denominator quietly ends your trend line. If the set must change, restate the baseline and say so in the report.
Why do AI answers change between runs, and how do you report that?
AI answers are probabilistic, so the same prompt can produce different results on different days, which is why AEO KPIs should be reported as a rate across a fixed prompt set over a period, never as a single snapshot.
The variance is measured, not anecdotal. SparkToro ran 12 identical prompts 2,961 times in late 2025 and found under a 1 in 100 chance that two runs return the same brand list (SparkToro). We ran our own test on 20 August 2026: one buyer prompt, 40 fresh-session runs across ChatGPT and Gemini, and no two runs produced the same brand list (full figures here). Semrush’s clickstream work adds a structural reason: ChatGPT switched web search on for 34.5% of queries in February 2026 and answered the rest from memory (Semrush), and the two routes produce different answers.
The reporting consequence: a one-week dip is weather. Report each KPI monthly, as a rate across the whole frozen set, and only treat a move as real when it survives two consecutive periods or the citations underneath the answers visibly changed. Write that rule into the client’s first report so the panic call never happens.
Our own numbers, August 2026
From the live weekly scans of the 3 UK B2B sites AEOscore monitors, measured 21 August 2026: across 551 answers that named a monitored brand, 22% carried no link to the brand’s domain. On ChatGPT alone the figure was 46%, against 8% on Gemini: measure citations alone and you miss almost half of ChatGPT’s recommendations. Across the same scans, the median non-brand mention rate per site is 6.2%, with the spread running from 0.4% to 22.9%: that is the realistic starting position for a B2B brand that has never done AEO work.
“Never report a single run. One AI answer is a coin flip, but ask forty and you have a measurement.”
What you send
A monthly client report template you can copy
How often should agencies report AEO performance to clients?
Report monthly with a quarterly deep-dive, because AI answers vary run to run and weekly reporting invites clients to react to noise rather than trend.
What should an agency put on the first slide of an AEO client report?
Put visibility (mention rate and share of voice against named competitors), the direction of travel versus last month, and the number of identified gaps fixed this month, because clients need to see both a score and evidence of work done.
Three views cover every client conversation, in this order:
Executive snapshot (one page)
Mention rate and share of voice against named competitors, direction of travel versus last month, gaps fixed this month, and one quoted answer that changed. Nothing else. This is the page the client forwards to their board.
Trend over time
Each KPI as a monthly line per engine, with the fix log plotted on the same axis so movement can be read against work shipped. Twelve months of history once you have it; the frozen prompt set is what makes this page honest.
Competitive standing by topic
Share of voice and recommendation position broken out by prompt topic, one row per topic, client versus the two nearest rivals. This page finds the topics where the client is losing and turns them into next month's fix list.
What KPI targets should you set for a new AEO client?
Set the baseline from the first month of scans, target the named competitor directly above the client in share of voice, and agree a monthly number of shipped fixes, because those three commitments are measurable and honest where an invented industry benchmark is neither.
Free template
KPI definitions, this-month and last-month columns, per-engine lines, competitor share of voice rows and a fix log, matching the structure on this page. Open it in Excel or Google Sheets and rename the columns for your client.
The other half
What to do when a KPI moves the wrong way
Most trackers stop at the number. The number only exists to trigger the next fix, so here is each wrong-way move mapped to its usual cause and the first thing to do:
Mention rate falls
Likely cause: A model or retrieval change on the engine's side, or a rival newly present in the sources the engine reads. Rarely something you broke this month.
First fix: Check the run log against engine release dates first, then read the citations on the changed answers. If a new source is doing the recommending, get the client onto that source.
Citation rate falls while mentions hold
Likely cause: The engine still knows the brand but has found better-structured pages to evidence with, or the client's pages have become unreadable to crawlers.
First fix: Re-check robots.txt and llms.txt allow the AI crawlers, then compare the winning cited pages against the client's: usually the rival page answers the question in its first paragraph and the client's does not.
Share of voice falls
Likely cause: A competitor shipped something the engines picked up: a roundup placement, a comparison page, a burst of reviews. Your client's mentions can hold steady while their share still falls.
First fix: Read the answers that changed and name the rival taking the share. The fix list is usually their citation list: pitch the roundup, publish the comparison, earn the reviews.
Non-brand inclusion stalls at zero
Likely cause: The engines have no page anywhere that connects the client to the buyer question, so there is nothing to retrieve and nothing to remember.
First fix: Publish a direct answer to each stalled prompt on the client's site (question as heading, answer in the first sentence), then get the client named in one third-party source the engines already cite for that question.
Recommendation position slips
Likely cause: The engines cite comparison and best-of pages in order, and the client has slipped down the underlying lists, or a rival has earned a stronger differentiator sentence.
First fix: Give the engines a reason to put the client first: a specific, dated claim on the client's own pages, repeated by a third party. Vague superlatives do not travel; numbers do.
Sentiment turns negative
Likely cause: The engine is quoting a real complaint (a review thread, a forum post) or misreading old pricing and stock as current.
First fix: Find the quoted source, fix what is fixable at the source, and publish the correction plainly on the client's site. Then re-scan until the answer changes; sentiment moves when its evidence moves.
Accuracy errors appear
Likely cause: Stale facts on pages the engines read: an old price on a forgotten landing page, a dead product in a directory listing, a namesake business being blended in.
First fix: Correct the stated fact everywhere it appears in visible body text, on the client's site first and the third-party listings second. Plain, current, visible text is what fixes stale claims.
AI referral traffic dips
Likely cause: Engine-side link behaviour changes month to month (ChatGPT's link-out rate has moved severalfold within a year), so the line moves for reasons that have nothing to do with visibility.
First fix: Check visibility KPIs before reacting: if mentions and citations held, the dip is plumbing, not presence. Note the engine-side change in the report and keep the trend line.
Closure rate falls
Likely cause: The fix list has outgrown the people fixing it, or gaps are being identified faster than anyone agreed to resolve them.
First fix: Cut the list, not the standard: agree a monthly fix budget with the client (a number of shipped fixes), rank gaps by the revenue closeness of their prompts, and carry the rest visibly rather than silently.
“A visibility report without a fix log is a weather report. Clients pay for the umbrella, not the forecast.”
What it costs
Pricing and what a fix cycle costs
Who is responsible for fixing the gaps a tracker finds?
Most AI visibility tools only report the gap, so the agency or an in-house technical team has to do the fixing: AEOscore is built to do both, pairing the weekly scan with a written fix plan, and on the £1,495 a month tier a named strategist and developer day that makes the changes.
AEOscore runs the whole loop described on this page as software with a person attached: the frozen prompt panel, weekly scans across all six engines, every KPI above scored per engine, a written fix plan for every gap, and a monthly report built to be forwarded. Built and run in the UK by Goreblimey Ltd, Cheltenham.
Pricing in one line: £395 one-off AI Visibility Audit, credited against your first month, £99 a month Monitor, £1,495 a month Monitor + Strategist. Prices exclude VAT. The AI Visibility Audit is £395 one-off: Full six-engine scan, fix plan, off-site targets, tone read, 30-minute call, PDF. Credited against your first month of either tier. Agencies: from £499 a month for 10 sites with a branded monthly report. Details on the pricing page, the scoring maths on the methodology page, and the difference between monitoring and fixing, spelled out, on monitoring vs fixing.
FAQ
AEO KPIs and client reporting: common questions
- Can an agency white-label AEO reporting for its clients?
- Yes: AEOscore's agency arrangement starts at £499 a month for 10 sites with a branded monthly report, so the agency rebills reporting under its own name while the scans, scoring and fix plans run underneath.
- Should we stop sending the keyword rankings report entirely?
- Keep rankings as an appendix for the commercial queries that still convert through clicks, and move the front of the report to AEO KPIs, because that ordering matches where the client's buyers now form shortlists.
- What does a realistic first month of AEO KPIs look like?
- Expect a non-brand mention rate in single digits (the median across the three UK B2B sites we scan is 6%), visibility on one or two engines out of six, and a citation rate around half the mention rate.
- How long do AEO fixes take to show up in the KPIs?
- Expect 4 to 12 weeks and uneven movement across engines: Perplexity and Google AI Overviews pick up new, well-sourced pages fastest, while ChatGPT and Claude lag because their answers depend partly on training data and on how often the page is referenced elsewhere.
- Is schema markup one of the KPIs or levers?
- No: Article and FAQPage schema are basic hygiene that help a page be read and indexed cleanly, and the content and sourcing are what earn citations, so add the markup but report content and off-site work as the levers.
- Do Google AI Overviews belong in the same client report?
- Yes, as one engine line among six, and note that Search Console folds AI Overview clicks into the Web search type rather than separating them, which is another reason the prompt-panel KPIs carry the reporting weight.
- What sample size makes an AEO KPI defensible in front of a client?
- Thirty prompts, three runs each, per engine, is the working minimum: it turns a coin-flip single answer into a few hundred scored answers a month, which is enough for the month-on-month trend to mean something.
Sources
Numbered sources
- Search Engine Land, reporting SparkToro and Similarweb clickstream data, Google zero-click searches reach 68% in early 2026: study. Accessed 21 August 2026. 68% of Google searches ended without a click to the open web, January to April 2026, US clickstream panel.
- Seer Interactive, AIO impact on Google CTR: September 2025 update. Accessed 21 August 2026. Organic CTR on queries showing an AI Overview fell to 0.61% by September 2025 from 1.76% in June 2024, a decline of around 61%. 3,119 informational terms, 42 organisations. Published November 2025.
- Pew Research Center, Google users are less likely to click on links when an AI summary appears in the results. Accessed 21 August 2026. Users clicked a traditional result on 8% of visits with an AI summary against 15% without. 900 US adults, 68,879 searches, March 2025 browsing data. Published July 2025.
- Forrester (John Buten), B2B Buyers Make Zero-Click Buying Number One. Accessed 21 August 2026. 94% of business buyers report using AI in their buying process (Forrester Buyers' Journey Survey, 2025). Published 22 January 2026.
- Demand Gen Report, reporting Labs by Demandbase, Demandbase: ChatGPT Referrals to B2B Websites Nearly Quadrupled in a Year. Accessed 21 August 2026. Monthly ChatGPT-referred visits to tracked B2B web properties rose to 2.6 million in June 2026 from roughly 645,000 in June 2025, a 303% increase. Demandbase customer platform data. Published 12 August 2026.
- Search Engine Journal, reporting Similarweb's 2026 Generative AI Landscape report, ChatGPT Links Out Most On Travel Queries, Data Shows. Accessed 21 August 2026. 6.8% of ChatGPT answers included outbound links as of May 2026 (up from 1.3% in June 2025). US desktop usage; Similarweb describes its figures as estimates. Published 27 July 2026.
- SparkToro (Rand Fishkin), AIs are highly inconsistent when recommending brands or products. Accessed 21 August 2026. 2,961 runs of 12 identical prompts: under a 1 in 100 chance that two runs return the same brand list, while aggregate frequency across many runs is far more stable. Published 28 January 2026.
- Semrush, ChatGPT traffic analysis: insights from 17 months of clickstream data. Accessed 21 August 2026. ChatGPT enabled web search on 34.5% of queries as of February 2026, with month-to-month swings between 15% and 66.3%. US clickstream panel. Published 7 April 2026.
- Google Analytics Help, Default channel group (includes the AI Assistants channel). Accessed 21 August 2026. GA4's default channel group added an AI Assistants channel in May 2026; custom channel groups apply retroactively where the default is reported not to backfill.
- AEOscore, First-party figures from live weekly scans of three UK B2B sites. Accessed 21 August 2026. Method described on this page; measured 21 August 2026 from the AEOscore scan database. The 40-run variance test is documented at /guides/ai-share-of-voice#variance.
What changed on this page
- : Guide published: the nine KPIs with formulas and starting ranges, the SEO-to-AEO comparison table, sampling method, monthly report template with CSV download, wrong-way playbook, and first-party figures from our live scans.
Start the client conversation with a real number
The free AEO score measures a domain across all six engines, human-checked, within 1 to 2 working days. Domain and email, no card.
- ChatGPT
- Claude
- Gemini
- Perplexity
- Google AI Overviews
- Google AI Mode