Evidence-first field guide · Updated September 2026

What is a good AI visibility score?

There is no defensible universal percentage. A useful benchmark begins with a frozen question set, explicit outcome rules, repeated observations and an uncertainty range.

9-minute readNo invented industry averagesWorked measurement example
The short answerA 30% mention rate can be strong, weak or meaningless depending on the questions, engines, markets, run count and definition of a “mention.” Compare your brand against a repeatable baseline built with the same protocol—not an unsupported cross-company average.

Why one benchmark number breaks down

Generative answers are not fixed search results. Repeating the same question can produce different brands and sources. Recent measurement research found substantial variation across repeated samples and warns that single-run point estimates can look more precise than the evidence allows.

Comparisons also fail when vendors count different outcomes. One report may count any brand name, another only linked citations, and another a top-three recommendation. A percentage without its denominator and collection protocol is not a benchmark.

Separate the metrics before comparing them

Mention rate

Answers naming the brand ÷ valid observations. Useful for unaided presence, but it does not show whether the statement is positive, accurate or linked.

Citation rate

Answers linking to the brand’s domain ÷ valid observations. This is narrower than a mention and depends on whether the experience displays sources.

Recommendation rate

Answers placing the brand inside a defined recommendation set ÷ relevant observations. Freeze what qualifies before collection.

Commercial outcome

Qualified visits, leads or purchases attributable to the channel. Visibility is a leading indicator—not revenue itself.

A defensible benchmark has six coordinates

CoordinateRecord it explicitlyWhy it changes the result
Audience and marketBuyer role, country and languageAvailable brands and evidence differ by market.
Question panelExact wording and buyer stageDiscovery, comparison and purchase questions produce different brand sets.
Engine conditionsProduct, model where visible, account state and dateRetrieval and generated responses vary across systems and time.
Outcome ruleMention, citation or recommendation definitionDifferent rules create different numerators.
Sample designQuestions × engines × repeated runsSmall or clustered samples create unstable point estimates.
UncertaintyInterval method and confidence levelA range communicates sampling uncertainty that a lone percentage hides.

Build your first baseline

  1. Define the decision. State what action the measurement could change: source repair, product-page clarification, comparison content or continued observation.
  2. Freeze a balanced question panel. Keep discovery, comparison, trust, risk and purchase questions separate so one stage cannot dominate the score.
  3. Write the counting rule first. Decide whether spelling variants count, whether citations must resolve, and how refusals or invalid answers enter the denominator.
  4. Repeat the observations. Record every run with its date and conditions. Do not keep only the most favorable answer.
  5. Report the count and range. Publish the numerator, denominator, point estimate and uncertainty interval together.
  6. Compare like with like. Rerun the same frozen panel before interpreting change. Treat a protocol change as a new baseline.

Worked example

Suppose 20 questions are tested across four engines with three runs each: 240 observations. If the brand appears 60 times, the observed mention rate is 25.0%. A 95% Wilson interval is approximately 19.9%–30.8%. Report all three values. A competitor’s isolated 29% point estimate is not enough to establish a meaningful difference without its own sample and uncertainty.

Use platform data without pretending it is the same metric

Google says eligibility for its generative search features still depends on ordinary indexing and Search fundamentals; it also provides generative-AI performance reporting in Search Console. Bing’s AI Performance report exposes citation totals and average cited pages across supported experiences. These first-party measures are valuable, but they describe platform activity and should not be silently merged with your controlled prompt-panel rate.

A practical dashboard therefore keeps three layers separate: controlled repeated observations, first-party platform visibility, and commercial outcomes. The goal is not a bigger vanity score. It is enough evidence to choose the next action.

Sources and scope

Limit: This guide provides a measurement framework, not a universal industry benchmark, ranking guarantee or statistical consulting service. Correlated observations may require cluster-aware or bootstrap analysis.

Turn the framework into a measured baseline.

Plan your sample with the free browser-local calculator, then use the Audit Kit when you need the complete evidence and client-delivery workflow.