Why one benchmark number breaks down
Generative answers are not fixed search results. Repeating the same question can produce different brands and sources. Recent measurement research found substantial variation across repeated samples and warns that single-run point estimates can look more precise than the evidence allows.
Comparisons also fail when vendors count different outcomes. One report may count any brand name, another only linked citations, and another a top-three recommendation. A percentage without its denominator and collection protocol is not a benchmark.
Separate the metrics before comparing them
Mention rate
Answers naming the brand ÷ valid observations. Useful for unaided presence, but it does not show whether the statement is positive, accurate or linked.
Citation rate
Answers linking to the brand’s domain ÷ valid observations. This is narrower than a mention and depends on whether the experience displays sources.
Recommendation rate
Answers placing the brand inside a defined recommendation set ÷ relevant observations. Freeze what qualifies before collection.
Commercial outcome
Qualified visits, leads or purchases attributable to the channel. Visibility is a leading indicator—not revenue itself.
A defensible benchmark has six coordinates
| Coordinate | Record it explicitly | Why it changes the result |
|---|---|---|
| Audience and market | Buyer role, country and language | Available brands and evidence differ by market. |
| Question panel | Exact wording and buyer stage | Discovery, comparison and purchase questions produce different brand sets. |
| Engine conditions | Product, model where visible, account state and date | Retrieval and generated responses vary across systems and time. |
| Outcome rule | Mention, citation or recommendation definition | Different rules create different numerators. |
| Sample design | Questions × engines × repeated runs | Small or clustered samples create unstable point estimates. |
| Uncertainty | Interval method and confidence level | A range communicates sampling uncertainty that a lone percentage hides. |
Build your first baseline
- Define the decision. State what action the measurement could change: source repair, product-page clarification, comparison content or continued observation.
- Freeze a balanced question panel. Keep discovery, comparison, trust, risk and purchase questions separate so one stage cannot dominate the score.
- Write the counting rule first. Decide whether spelling variants count, whether citations must resolve, and how refusals or invalid answers enter the denominator.
- Repeat the observations. Record every run with its date and conditions. Do not keep only the most favorable answer.
- Report the count and range. Publish the numerator, denominator, point estimate and uncertainty interval together.
- Compare like with like. Rerun the same frozen panel before interpreting change. Treat a protocol change as a new baseline.
Worked example
Suppose 20 questions are tested across four engines with three runs each: 240 observations. If the brand appears 60 times, the observed mention rate is 25.0%. A 95% Wilson interval is approximately 19.9%–30.8%. Report all three values. A competitor’s isolated 29% point estimate is not enough to establish a meaningful difference without its own sample and uncertainty.
Use platform data without pretending it is the same metric
Google says eligibility for its generative search features still depends on ordinary indexing and Search fundamentals; it also provides generative-AI performance reporting in Search Console. Bing’s AI Performance report exposes citation totals and average cited pages across supported experiences. These first-party measures are valuable, but they describe platform activity and should not be silently merged with your controlled prompt-panel rate.
A practical dashboard therefore keeps three layers separate: controlled repeated observations, first-party platform visibility, and commercial outcomes. The goal is not a bigger vanity score. It is enough evidence to choose the next action.
Sources and scope
- Google Search Central: optimizing for generative AI features — indexing, people-first content and Search Console measurement guidance.
- Google Search Central: generative AI performance reports — first-party reporting scope.
- Bing Webmaster Blog: AI Performance — definitions for citations and average cited pages.
- Quantifying Uncertainty in AI Visibility — repeated sampling, variability and interval reporting.
Limit: This guide provides a measurement framework, not a universal industry benchmark, ranking guarantee or statistical consulting service. Correlated observations may require cluster-aware or bootstrap analysis.
Turn the framework into a measured baseline.
Plan your sample with the free browser-local calculator, then use the Audit Kit when you need the complete evidence and client-delivery workflow.