Skip to content
Learn

What KPIs Should You Track for Generative Engine Optimization?

The KPI set for generative engine optimization: formulas, denominators, GA4 setup, per-engine measurability, benchmarks, and a one-page report template.

Line icon of an eye above a two-by-two grid of tiles marked with a check mark, an X, a triangle, and a circle, connected by dotted lines on a green background.

Most generative engine optimization measurement dashboards default to reporting simple ranking vs competitors. But, a rank tracker may show position three for the category term while an AI answer to the same question phrased differently recommends two competitors instead. You need to know what the engines are looking for before you can reliably benchmark.

Generative engine optimization (GEO) KPIs should measure brand presence and performance inside AI-generated answers, meaning whether you appear, how you're described, and what those appearances produce.

Here's the set of measures we'd put on the report, with the formula behind each line.

What are GEO KPIs?

Generative engine optimization KPIs record whether an AI engine names your brand, cites your pages, describes you accurately, and sends visitors who convert.

We use the term AI visibility for that whole set, with GEO and AEO (answer engine optimization) as the narrower labels most vendors sell under. The working set has seven parts:

  • Citation and brand mention frequency
  • AI answer inclusion rate
  • AI share of voice
  • AI referral traffic
  • Brand sentiment in AI answers
  • Entity signals
  • AI-sourced conversions, leads and revenue

A mention is your brand named in the answer text, and a citation is your URL used as a source for that answer. In a 2026 index built from 126 million real user prompts across 22 industries, fewer than 25% of the most-mentioned brands were also among the most-sourced. A brand can be recommended constantly and cited almost never, so the two numbers need separate lines.

A third line is recommendation, whether the answer presents you as a fit for the buyer's situation. Inclusion counts a mention that rules you out as a win, so log recommended, listed, and mentioned with a caveat as separate states.

Entity signals, whether engines associate your brand with the right category, competitors, and attributes, are the least standardized of the seven. No vendor publishes a formula, so most teams fold them into the accuracy review.

How to calculate each KPI

Write down the numerator, denominator, question panel, and refresh cadence for each KPI before the first report ships. Vendors define the same word differently, and we've watched a mid-year tool switch get read as a visibility collapse.

Before any formula, list the buying questions your buyers need answered (comparisons, alternatives, pricing, fit) and rank them by measured demand. Each question gets a few phrasings, which is where prompts belong: as tests of a question you already know matters. Keep questions that name your brand out of the inclusion denominator and read them separately, because they measure what the engine says about you, not whether it volunteers you. Start small and expand until every competitor you report on appears in some answers.

Citation and brand mention frequency

Mention frequency is sampled responses that name your brand ÷ total sampled responses. Citation frequency uses the same denominator with a different numerator, responses that use at least one URL from your domain as a source. Run both per engine. Sample daily and report a seven-day rolling figure, because answers to the same question change between runs. In our own CheckThat data, among page-and-question pairs observed on three or more days, 59% held a citation on exactly one day and the average cited page held it on 30% of the days observed. That churn runs in both directions, so a single run is not a reading.

AI answer inclusion rate

Inclusion rate is questions with at least one brand appearance ÷ questions on the panel × 100. It's the headline number for a leadership slide because it answers the first question executives ask, which is whether you show up at all.

Some tools count per response rather than per question, and some divide only by responses that named a brand. Pick one denominator, write it in the report footer, and keep it for the year.

AI share of voice

Share of voice is your brand's mentions ÷ mentions of all tracked brands across the same question panel, × 100. Fix the competitor list in writing, because adding a competitor mid-quarter shrinks everyone's share and looks like a loss you didn't take.

Report it per engine. The same 2026 index found ChatGPT and Google AI Mode agreed on which brands to mention 67% of the time and on which sources to cite only 30% of the time, so an average hides real differences. Use it for direction, never as an audited market share.

AI referral traffic

AI referral sessions are GA4 sessions whose source matches a known AI hostname. AI referral share is AI sessions ÷ total sessions × 100. Use total sessions as the denominator and say so, because published studies use three different denominators and can't be compared.

Treat the result as a floor, because clicks from AI mobile apps often arrive without a referrer header and land in Direct, and GA4 classifies Google AI Overviews and AI Mode as google / organic.

Brand sentiment in AI answers

Sentiment classifies each answer that names your brand as positive, negative, or neutral. We keep the scheme coarse on purpose so a human can audit it. Net sentiment collapses the mix to one number, (positive − negative) ÷ total × 100, running from −100 to +100.

Sentiment is a quality KPI, so read it against inclusion rate. High inclusion with a negative mix means engines are recommending against you at scale.

AI-sourced conversions, leads and revenue

AI-sourced conversion rate is key events (form fills, trial starts, demo requests) from the AI channel ÷ AI channel sessions. Lead and pipeline counts come from the CRM: contacts whose original traffic source is an AI referrer, the opportunities they open, and the pipeline value attached. Report them beside the organic equivalents for the same period.

Leadership judges the program on this line, and it has the weakest data of the six with a formula. Volumes are small, referrer stripping removes an unknown share, and B2B sales cycles push this quarter's citations into next quarter's pipeline. Label the tracked figure as a floor.

What carries over from SEO reporting?

Rank and click-through rate describe a results page, and an AI answer has neither a position number nor a reliable click. Each SEO line gets a generative engine optimization counterpart that is per-engine, measured against a fixed question panel and competitor set, and never promises a click:

SEO KPIGEO counterpart
Keyword rankAI answer inclusion rate
Organic CTRCitation frequency
Organic sessionsAI referral sessions
BacklinksOff-site mentions and citations
Keyword volumeQuestion demand
Position per keywordShare of voice per question panel

Across 3,119 search terms and 25.1 million organic impressions, organic CTR on queries where an AI Overview appeared fell from 1.76% in June 2024 to 0.61% in September 2025, and paid CTR on the same queries fell from 19.70% to 6.34%. A CTR line on its own no longer tells you what happened.

Does SEO still matter in 2026? Yes. Engines using live retrieval need pages they can crawl and index, so search fundamentals feed citation even though a top Google ranking doesn't guarantee AI placement. Keep the rank report and add the GEO block beside it.

Benchmarks for each KPI

No published study gives a citation or inclusion rate with disclosed methodology that you can adopt as a target, so your benchmark is your own baseline against named competitors on the same question panel. CheckThat, which powers AI visibility in the GrowthX platform, supplies the category context for that baseline. As of September 2026 it measures more than 5,800 brands and more than 2.6 million AI responses across nearly 200 B2B software categories, which tells you whether a 12% inclusion rate is normal for your category.

For traffic and conversion, the population behind each figure matters more than the number:

KPIPublished figurePopulation and period
AI referral share of sessions0.5%97 B2B websites, 28.9M sessions, July 2025 to June 2026
AI referral conversion rate4.87% vs 4.60% for organic, not statistically significant54 websites, six months each
Conversion rate by engineChatGPT 15.9%, Perplexity 10.5%, Claude 5.0%, Gemini 3.0%, Google organic 1.76%One B2B client, about 11,000 AI sessions

The two conversion rows describe different populations. A 54-site average with a strict conversion definition lands at parity with organic, while one client with a favorable question mix shows ChatGPT converting at nine times Google organic. The 0.5% traffic figure carries its authors' own caveat, which applies to every row: GA4 can't report all AI traffic separately.

No cross-vendor benchmark exists for sentiment or accuracy. For the first two quarters, set every target as a delta from your own baseline, because other teams measured the published numbers on their own question panels.

Setting up GA4 tracking

Google added a native AI Assistant default channel to GA4 on May 13, 2026. It assigns medium = ai-assistant when it detects a recognized referrer from ChatGPT, Gemini, Deepseek, Copilot, or Grok, and its channel definitions leave Google AI Overviews and AI Mode in Organic Search. Perplexity isn't on the list and lands in Referral. Treat the channel as the base layer and add two things on top.

First, a custom channel group. In Admin > Data display > Channel groups, create a channel named AI Assistants with the condition Source matches regex, then reorder it above Referral so it fires first. This pattern, published in July 2026, covers the hostnames most B2B teams need:

^(chatgpt\.com|chat\.openai\.com|claude\.ai|perplexity\.ai|www\.perplexity\.ai|gemini\.google\.com|bard\.google\.com|copilot\.microsoft\.com|deepseek\.com|grok\.com|x\.ai|meta\.ai|you\.com|poe\.com)$

Don't add bare google.com or bing.com, because that would bucket every Google Search session into the AI channel.

Second, Search Console's Generative AI performance report, worldwide since August 31, 2026, gives impressions for AI Overviews and AI Mode. It's the only official feed for the surfaces GA4 can't separate.

For links you control, set utm_source to the engine name, utm_medium to exactly ai-assistant so GA4's native channel recognizes it, and utm_campaign to a topic slug. Never publish a canonical URL carrying those parameters, because a tagged canonical mislabels every visitor who copies it.

Measuring visibility by engine

ChatGPT, Perplexity, Gemini, and Claude each offer an official API you can sample, and none of the four documents its API output as identical to the consumer app. OpenAI's web search documentation says the API's web search doesn't exactly match consumer ChatGPT. Every API-based figure is therefore a proxy, and tools that capture the consumer interface instead do so against each vendor's terms of service. Google AI Overviews has no public API, so its data comes from third-party SERP capture.

Referral traffic and conversions are directly measurable for ChatGPT and Perplexity on the web, partial for Gemini and Claude because some sessions arrive with no referrer, and unavailable for AI Overviews because it merges into google / organic.

Accuracy and hallucination rate

Hallucination rate is answers with an incorrect brand claim ÷ answers naming the brand × 100. It decides whether a high inclusion rate is an asset or a liability. One vendor's snapshot of more than 158,000 claims put answer-engine inaccuracy rates between 4.3% and 6.8%, which is low enough to ignore until the wrong fact is your pricing.

Keep one canonical page of brand facts (pricing model, plan names, integrations, headquarters), and for each sampled answer have a reviewer classify every brand claim as accurate, inaccurate, or unverifiable, assigning the sentiment tier in the same pass. Read the negatives individually, because a "negative" answer is often an accurate description of a limitation you have, and no content fix removes that.

The remediation loop runs on the citation data you already collect:

  • Find the source. Use the cited URLs in the answer to find which page the engine relied on.
  • Fix or counter it. Correct your own page, or publish a page on your domain that states the fact plainly.
  • Re-sample the same questions. Track the rate on the fixed panel and record the date of each fix beside it. The dated pair shows what changed and when, not that the fix caused it.
  • Watch 404s in the AI channel. Engines cite URLs that no longer exist, and a redirect recovers the visit and the citation.

Attributing AI referral sessions to pipeline

Nothing joins a citation to a deal. What can be joined is a referred session, which is a subset of the buyers who saw you in an answer. A leadership report on AI-sourced pipeline should present a range, with a tracked floor from referrer-based sessions and a modeled ceiling from self-reported source.

The tracked floor comes from GA4 key events on the AI channel joined to CRM contacts by original source. HubSpot has carried a distinct AI Referrals traffic source since February 2025 that recognizes nine platforms and flows into its attribution reports. Other CRMs need a custom property populated from the GA4 channel or a hidden form field.

The ceiling comes from a "how did you hear about us" field with the engines listed by name. It captures the buyer who saw you in ChatGPT and Googled your name a day later, which referrer data records as organic or direct. The gap between the two is your own dark traffic figure, not an industry estimate.

The GrowthX platform does not report revenue or dollar-value attribution, and no visibility tool should. That figure belongs in the CRM with the deal data.

Tools that report GEO KPIs

Five tools cover most of the generative engine optimization KPI set. Profound, Peec AI, Otterly.ai, Semrush's AI Visibility Toolkit, and HubSpot AEO all report a visibility or inclusion score, share of voice, and sentiment, and all but Peec add a citation view. Profound ties AI-referred conversions in through Google Analytics, and engine coverage runs from HubSpot's three (still labeled beta) to Profound's nine on its enterprise tier.

HubSpot AEO's visibility metrics aren't documented as feeding contact or deal attribution, so its AI Referrals source is still the piece that reaches the pipeline. Any tracker that fetches your pages live can become part of the signal it measures.

AI visibility in the GrowthX platform is powered by CheckThat, whose measurement model has four parts: Presence (whether AI recommends you when buyers evaluate), Reputation (what the outside world says about you), Perception (what story AI tells about you and whether it is accurate), and Influence (how much control you have over the sources behind the answer). Inclusion and citation feed Presence, sentiment and accuracy sit under Perception, and cited-source data feeds Influence. The model is still being built out, so treat the four names as vocabulary rather than four numbers on a report today.

The answer side appears in the platform's AI Engine Visibility report: citation status, a visibility score, sentiment and trend, share of voice, top-3 rate, per-engine rankings, cited-domain leaderboards, and the buyer questions where you win and lose, across five engines including Google AI Overviews and Google AI Mode. The referral side, sessions from each answer engine against the prior period, sits in the Progress report, the same split this article keeps. Both sit beside page and portfolio reporting in one workspace, so a gap can go straight into a brief. The platform is built to run a whole content program, so if monitoring is the only job, one of the trackers above is the right buy.

A one-page report and review cadence

The monthly report fits on one page when every KPI carries its formula, data source, and refresh cadence in one row, and the footer states the question panel and its phrasings, the competitor set, and the denominators in use:

KPIData sourceRefresh
Inclusion rateQuestion panel, per engineDaily sample, weekly rollup
Mention and citation frequencyQuestion panel, per engineDaily sample, weekly rollup
Share of voiceQuestion panel, fixed competitor setWeekly rollup
Sentiment mixQuestion panel, human review of negativesScored each run, reviewed monthly
Hallucination rateHuman review against the fact sheetMonthly
AI referral shareGA4 custom channelWeekly volume, monthly share
AI-sourced conversion rateGA4 key eventsMonthly
AI-sourced leads and pipelineCRM plus self-reported sourceMonthly, pipeline quarterly

A weekly swing should hold across two consecutive rollups before anyone credits it to the team's work, since a platform update can move a score as far as a content change can. The monthly one-pager goes to leadership with the deltas and the hallucination fix list. Each quarter, revise the question panel and competitor set and log every change, because each one resets the baseline.

Build the question panel this month, in a spreadsheet if that's what you've got, and get the GA4 channel group live before the next reporting period. If maintaining this across separate tools is the part that stalls, book a demo to see how the GrowthX platform runs it in one workspace.