Measure AI search visibility across repeated answers under recorded conditions. A single response shows one outcome. It does not establish how often customers will see the same brands, sources, or claims.
Define the metrics
Define each metric and its denominator before collecting results. The measures below answer different questions.
- Mention rate: the share of sampled answers in which your brand appears at all, in any role.
- Recommendation rate: the share in which the answer actually recommends you, rather than mentioning you in passing or unfavorably.
- Citation rate: the share in which the answer links to your site. Record third-party citations about the brand separately.
- Share of voice: your mentions as a fraction of all brand mentions across the sampled answers, the competitive version of mention rate.
- Prominence: where you appear when you appear: first pick, list position, or a trailing aside.
- Sentiment: how the answer characterizes you when it names you.
- Accuracy: whether what the answer says about you is true and current.
A brand can be mentioned without being recommended. A source can be cited in an answer that recommends a competitor. Keep these outcomes separate so the report remains useful for diagnosis.Research sources and limitations. Identifies the available report data and the analysis still needed to support earlier draft figures.
Sampling repeated answers
Repeat prompts and record variation. The measurement literature identifies changes across wording, runs, and time. Your report should state which of those variables were held constant.Don't Measure Once: Measuring Visibility in AI Search, the paper making the empirical case that one-off observations are unreliable and visibility should be reported as a distribution.
The following is a practical starting protocol. An API can expose useful settings and logs, but an API test is not necessarily equivalent to a consumer chat session.
- Runs: ten completed runs per prompt and condition can be a starting point. Increase the sample when the variation or decision requires more evidence.
- Search conditions: compare the same prompt with retrieval disabled and enabled where supported. Record actual tool use when available.
- Wording: include relevant paraphrases and track each version separately.
- Buyer requirements: include the customer differences that affect product suitability. Keep those results separate before combining them.
- Products: select the experiences your audience uses and record what the test actually measures.
- Versions: record a fixed model identifier or changing alias where exposed. An API alias does not necessarily reproduce the consumer product.
Report counts with rates, such as “recommended in 7 of 10 completed answers.” Small samples have substantial uncertainty. A rate is comparable over time only when the prompts, scoring rules, and relevant conditions are comparable.
The confounders
Record model version where exposed, product, account state, tools, location, prior conversation, and supported generation settings. Mark unknown values explicitly. Search availability and actual search execution are separate facts.
Reading the split in your data
Compare search-enabled and search-disabled results to identify differences associated with the test conditions. A difference suggests where to investigate; it does not prove that one answer came exclusively from training knowledge or that ranking caused an omission.Research sources and limitations. Identifies the available report data and the analysis still needed to support earlier draft figures.
Evaluating a tracking method
Evaluate a tracking tool by the method it documents. Ask which user experience it measures, how many responses it collects, and which settings it can identify. Choose daily, weekly, or monthly sampling according to source volatility and the decisions the report supports.
Require the prompt set, sample counts, scoring definitions, collection failures, and available configuration records. Review actual answers as well as aggregate scores. Building a Visibility Tracking System covers implementation.
Use first-party reporting where it exists
Google’s Generative AI performance report provides impressions for AI Overviews and AI Mode. Its documentation states worldwide rollout on August 31, 2026. The report can group results by dimensions such as page, country, and device.Google Generative AI performance report. Describes impression reporting, aggregation, and exports. Read unavailable values carefully: the export can represent them as zero.
Keep that report separate from sampled brand recommendations. A link impression is a provider-recorded exposure. A mention rate is calculated from the answers you collected. Neither is automatically a count of customers who considered buying.
Save the reporting period, dimensions, and export with the analysis. A missing report or unavailable value must not become evidence of zero visibility. For referral and conversion reporting, use Attribution and Business Impact.
Score the same answer the same way
Write a short scoring rule before reviewing answers. In the worksheets supplied with this guide, each brand receives at most one mention, one recommendation, and one site-citation score per completed answer. Repeating a brand name five times does not create five observations.
| Answer text or condition | Mention | Recommendation | Site citation |
|---|---|---|---|
| “Example Cooling is a suitable option,” with a link to its site | 1 | 1 | 1 |
| “Example Cooling does not serve that area” | 1 | 0 | Score from the actual link |
| The answer links to a comparison page on the business’s site but recommends other companies | Score from the answer text | 0 | 1 |
| The business is absent from a completed answer | 0 | 0 | 0, unless its site is cited separately |
| The request fails before returning an answer | Not scored | Not scored | Not scored |
For this protocol, a site citation requires a link to the business’s domain. A linked third-party article about the business is recorded as a third-party citation. Keep aliases and domain mappings in the scoring notes so a rebrand or shortened name does not change the count accidentally.
Score accuracy at the claim level. An answer can state the correct service area and the wrong fee. “Two claims checked; one incorrect” is more informative than assigning one vague accuracy score to the whole answer.
Keep the denominator visible
Mention rate equals completed answers mentioning the brand divided by completed answers in the selected prompt group and condition. Recommendation and citation rates use the same denominator. Report the attempted count separately.
For share of voice, count distinct brand-answer pairs. If Brand A appears in 9 answers, Brand B in 12, and Brand C in 9, A has 9 / 30 = 30% of mentions across those three tracked brands. Adding another tracked competitor changes that denominator even when A’s answers do not change.
Do not average percentages from unequal samples without checking the weighting. Five mentions in ten answers and one mention in two answers both equal 50%, but they contain different amounts of evidence. Preserve the counts and report results by prompt before combining them.