Back to guides
AI Visibility MeasurementUpdated July 22, 202613 min

AI Visibility Metrics: Mentions, Position, Share of Voice, Sentiment, and Citations

A defensible measurement framework for turning repeated AI answers into comparable visibility, accuracy, and source metrics.

Key Takeaways

  • Measure a controlled set of prompt runs, not a supposed universal AI ranking.
  • Keep mentions, recommendation position, sentiment, accuracy, and citations as separate signals.
  • Freeze the prompt set and execution conditions before treating week-to-week movement as a trend.

Start With The Observation, Not The Dashboard

AI visibility is not one number. A generated answer can name a brand without recommending it, recommend it without linking to it, cite a publisher that discusses it, or describe it inaccurately. Compressing all of those outcomes into a single score hides the part a team can actually improve.

The basic unit of analysis should be one observation: the exact prompt, platform, model or product mode, search setting, country, language, date, and run number paired with the complete answer and its visible sources. Two answers are comparable only when those conditions are materially the same. Save the raw response so every score can be audited later.

Generated answers vary between runs. A single screenshot is evidence that an answer occurred, not evidence of stable market visibility. For an exploratory baseline, run every core prompt at least three times per platform under documented conditions. Use more repetitions for important decisions, and report the denominator beside every percentage.

  • Observation ID: prompt version + platform + mode + locale + date + run number.
  • Valid run: the system returned a complete answer that addressed the task.
  • Invalid run: an error, refusal, empty response, or answer unrelated to the prompt; log it, but do not silently score it as an absence.
  • Comparison rule: do not mix different prompt versions, regions, or search-on and search-off modes in the same trend line.

Measure Mentions And Prompt Coverage Separately

A mention is a normalized reference to the tracked brand, product, or accepted alias in the answer text. Decide the aliases before scoring. Exclude a brand that appears only inside a source title, URL, user prompt, or boilerplate unless your stated method intentionally counts those locations.

Mention rate answers how often the brand appeared across repeated runs: valid runs mentioning the brand divided by all valid runs, multiplied by 100. If 43 of 120 valid runs mention the brand, the mention rate is 35.8%. Prompt coverage answers a different question: unique prompts that produced at least one mention divided by unique prompts tested. If 18 of 30 prompts produce a mention in any run, prompt coverage is 60%.

Use both. A brand may have broad but unstable coverage, appearing once across many prompts, or narrow but consistent coverage, appearing on every run for only a few prompts. Segment the results by buyer stage and topic so a high branded-prompt score cannot conceal weak visibility in unbranded recommendation prompts.

  • Mention rate = mentioning runs ÷ valid runs × 100.
  • Prompt coverage = prompts with at least one mention ÷ prompts tested × 100.
  • Unbranded mention rate should be reported separately from branded diagnostic prompts.
  • A mention is not automatically a recommendation, positive statement, or citation.

Score Recommendation Position Without Inventing Precision

Position is appropriate only when the answer presents an ordered recommendation set. Record 1 for the first recommended brand, 2 for the second, and so on. If the prose names several options without a meaningful order, label the position unordered rather than forcing a rank. If the brand is absent, record absent, not last place.

Top-three rate is the share of valid ordered recommendation runs in which the brand appears in positions one through three. Mean reciprocal rank gives earlier placements more weight: for each ordered recommendation run, score 1 divided by the position, and score 0 when absent; then average the scores. Positions 1, 3, absent, and 2 produce (1 + 0.33 + 0 + 0.5) ÷ 4 = 45.8%. Publish the plain-language top-three rate beside it so stakeholders can interpret the result.

Do not compare position from a 'best tools' list with a passing mention in explanatory prose. Define a qualifying recommendation prompt and a qualifying ordered answer before collection begins. This reduces the temptation to rescore ambiguous answers after seeing the result.

  • First-position rate = ordered recommendation runs placing the brand first ÷ valid ordered recommendation runs × 100.
  • Top-three rate = ordered recommendation runs placing the brand in positions 1–3 ÷ valid ordered recommendation runs × 100.
  • Reciprocal-rank score per run = 1 ÷ position; absent = 0; unordered = not applicable.
  • Keep ordered, unordered, and non-recommendation answers in separate fields.

Calculate Share Of Voice Against A Fixed Competitor Set

AI share of voice describes how much of the tracked recommendation set a brand occupies relative to named competitors. It is not market share and should not be presented as one. Freeze the competitor list for the reporting period, normalize product aliases, and decide whether parent brands and products are separate entities.

A simple prompt-set share of voice is the brand's qualifying recommendation appearances divided by qualifying appearances for every tracked brand, multiplied by 100. If the monitored brands receive 80 qualifying appearances and your brand receives 22, its share is 27.5%. Because one answer may recommend several brands, this denominator is total brand appearances, not total responses.

For a weighted view, assign each prompt a business weight before running it—for example, 3 for a high-value category decision, 2 for a comparison, and 1 for early research. Multiply each qualifying appearance by the prompt weight, then calculate the same share. Always publish the weights. Retrospectively increasing the weight of prompts where the brand performs well turns a measurement into marketing copy.

  • Unweighted share of voice = your qualifying appearances ÷ all tracked-brand qualifying appearances × 100.
  • Weighted share of voice = your weighted appearances ÷ all tracked-brand weighted appearances × 100.
  • Use an 'other brand' field to notice emerging competitors without changing the frozen comparison set mid-period.
  • Report share by topic as well as overall; an aggregate can hide a decisive category gap.

Separate Tone From Factual Accuracy

Sentiment asks how the answer frames the brand. A practical human-coded scale is positive, neutral or mixed, and negative. Store the supporting sentence with the label. Automated sentiment can help triage large samples, but product comparisons, caveats, negation, and domain language make unsupervised labels unreliable enough to require review.

Message accuracy asks whether verifiable statements about the brand are correct. Create a fact sheet containing current pricing model, audience, capabilities, availability, locations, and other claims that matter to buyers. Score each material statement as accurate, outdated, inaccurate, or unverifiable. The material inaccuracy rate is observations containing at least one material error divided by observations that mention the brand.

For example, if 43 observations mention the brand and three state an obsolete price or nonexistent feature, the material inaccuracy rate is 7.0%. That is more actionable than a generic negative-sentiment score: the team can trace repeated errors to pages and third-party sources that may need correction. Never label a fair limitation or unfavorable comparison as inaccurate merely because it is commercially inconvenient.

  • Tone label: positive, neutral or mixed, negative, or unclear.
  • Accuracy label: accurate, outdated, inaccurate, or unverifiable.
  • Material inaccuracy rate = mentioning observations with a material error ÷ mentioning observations × 100.
  • Quality control: have a second reviewer rescore a sample and resolve disagreements in the rubric.

Track Citations As A Source System

A brand mention and a source citation are different events. ChatGPT Search may show inline citations and a Sources panel, while Claude web-search answers include source links. A response might recommend a brand based on third-party coverage without citing the brand's site, or cite the brand's documentation without naming it in the recommendation. Capture visible source URLs independently from the answer text.

Define a citation opportunity as every valid run conducted in a web-enabled mode designed to display sources, whether or not that particular answer returns a visible source. Overall citation presence is observations with at least one visible source divided by citation opportunities. Owned citation rate is observations with at least one owned-domain citation divided by all citation opportunities. Earned citation rate uses observations with at least one independent source that substantively discusses the brand. Count each qualifying observation once in these run-level rates, even if it contains several qualifying links. Keep owned, earned, directory, forum, and irrelevant sources as separate types.

Source-gap analysis is often the most useful output. For each high-value prompt where a competitor is recommended and your brand is absent, list the domains cited in that answer. Repeated domains become research targets: understand what information they provide, whether your product genuinely qualifies, and whether you can contribute a useful, accurately disclosed source. A citation count alone does not prove that a source caused the answer.

  • Owned citation rate = observations with at least one owned-domain citation ÷ citation opportunities × 100.
  • Citation diversity = count of unique relevant referring domains in the reporting period.
  • Source concentration = relevant domain appearances from the top three domains ÷ all relevant domain appearances × 100; count each canonical domain at most once per observation.
  • Store the final URL, page title, source type, associated prompt, and the claim the source appears to support.

Build A Baseline You Can Defend

Freeze a core prompt set, competitor set, scoring rubric, and execution protocol for at least one reporting cycle. Run the same conditions on a consistent cadence, keep raw answers, and annotate model, interface, or product changes. Treat small movements as directional until they repeat across several collection periods; generated-answer samples are not a census of all users.

A useful scorecard shows denominators and absolute counts beside rates: 43 of 120, not only 35.8%. Split results by platform, prompt cluster, branded versus unbranded intent, and search mode before calculating an overall view. If you publish a composite score, document its components and weights and retain the component metrics; otherwise the score cannot explain what improved.

This framework is deliberately distinct from ordinary organic-search reporting. Search Console's dedicated generative AI report measures impressions for your URLs inside supported Google Search generative features. Controlled prompt monitoring measures what selected generated answers say. Neither should be substituted for the other, and neither alone proves revenue impact.

  • Baseline: at least three repeated runs per core prompt and platform for exploratory monitoring.
  • Trend rule: compare identical prompt versions and conditions, with changes annotated rather than hidden.
  • Zero-denominator rule: report N/A, not 0%, when a segment has no qualifying observations; retain the zero count so the missing basis is visible.
  • Decision rule: investigate clusters and raw answers before acting on an aggregate movement.
  • Outcome link: connect priority clusters to qualified visits, trials, pipeline, or another business measure where attribution is available.

Apply the measurement model

Benchmark the brands that appear in the same buyer answers

Move from definitions to a recurring competitor view with answer-level evidence behind every metric.

Explore competitor tracking