Back to templates
Reporting TemplateUpdated July 22, 202612 min

Weekly AI Visibility Report Template

A copy-ready weekly report for turning repeated AI answers into a transparent scorecard, source-gap review, and prioritized action list.

Key Takeaways

  • Lead with decisions and material changes, then show the denominators and raw evidence behind them.
  • Keep Search Console generative impressions separate from sampled prompt-monitoring rates.
  • End every report with a small, owned action list and a methods note that makes comparisons reproducible.

What This Weekly Report Is For

The report should answer three questions: what materially changed, what evidence explains the change, and what will the team do next? It is not a gallery of favorable screenshots. A useful report includes absences, factual problems, failed runs, source concentration, and method changes alongside positive results.

Use this template after you have a fixed buyer-prompt set and scoring rubric. It works for manual research or exported monitoring data. If the organization has access to Search Console's Generative AI performance report, include its impression data in a separate first-party panel rather than combining it with prompt-sample rates.

Choose a consistent comparison: week over week for an operational pulse and a four-week baseline for context. Generated answers can move between runs, so investigate repeated cluster-level changes before announcing a trend. For low-volume samples, show counts beside percentages and use 'directional' language.

Prepare The Data Before Writing Commentary

Complete the data-quality block first. Confirm that core prompt version, platform modes, locale, number of repetitions, competitor set, and collection window match the comparison period. If they do not, label the comparison as partial or create a new baseline. Never hide a prompt-set or model change in a footnote after presenting a large percentage movement.

Keep one row per observation: prompt ID, platform, mode, locale, run, answer, visible sources, mention label, recommendation position, competitor labels, tone, accuracy, and reviewer. Invalid runs stay in the log with a reason. Exclude them from valid-run rates only according to the predeclared method.

Deduplicate source URLs to their canonical form where practical, but retain the observed URL. Normalize brand aliases using the same dictionary every week. Spot-check the largest score changes and every material inaccuracy against the raw answer before finalizing the report.

  • Same core prompt versions and business weights as the comparison period.
  • Same platforms, known modes, locale, language, and search settings.
  • Expected run count reconciled with valid, invalid, and missing runs.
  • Brand aliases, competitors, source types, and scoring rubric unchanged or explicitly annotated.
  • Raw answers retained and a reviewer assigned to material findings.

Use A Small Set Of Reproducible Formulas

Mention rate equals valid runs mentioning the brand divided by valid runs. Top-three rate equals ordered recommendation runs placing the brand in positions one through three divided by valid ordered recommendation runs. Share of voice equals your qualifying brand appearances divided by qualifying appearances for all frozen tracked brands. Each formula needs its absolute numerator and denominator in the scorecard.

For sources, owned citation rate equals observations with at least one owned-domain citation divided by all citation opportunities. A citation opportunity is every valid answer in a web-enabled, source-visible mode, including answers that return no visible source. Count an observation once even when it contains several owned links. Material inaccuracy rate equals mentioning observations that contain at least one material factual error divided by all observations mentioning the brand.

Do not average platform percentages unless you have a declared weighting method. Report ChatGPT and Claude separately, then show an explicitly weighted combined view only if it serves a decision. Whenever a denominator is zero—no valid ordered recommendations, tracked-brand appearances, citation opportunities, or brand mentions—report the rate as N/A rather than 0%, and retain the zero count. Search Console generative impressions are first-party counts and remain outside these formulas.

  • Mention rate = mentioning valid runs ÷ valid runs × 100.
  • Top-three rate = top-three ordered recommendation runs ÷ valid ordered recommendation runs × 100.
  • Share of voice = your qualifying appearances ÷ all tracked-brand qualifying appearances × 100.
  • Owned citation rate = observations with at least one owned-domain citation ÷ all citation opportunities × 100.
  • Material inaccuracy rate = mentioning observations with a material error ÷ mentioning observations × 100.

Turn Changes Into Evidence-Based Findings

A finding should contain the observation, its scope, the evidence, and the plausible next check. 'Visibility improved 20%' is incomplete. 'Unbranded mention rate increased from 18 of 90 to 28 of 90 valid ChatGPT Search runs, concentrated in three security prompts; two newly recurring publisher pages were visible sources' is reviewable.

Separate observation from explanation. The recurring publisher pages may be relevant, but the sample does not prove that they caused the recommendations. Phrase the next step as a hypothesis: review those pages, compare their selection criteria with product evidence, and check whether the pattern persists in the next collection.

Include one or two raw-answer excerpts only when they clarify a material pattern, and link internally to the complete observation. Avoid cherry-picking a flattering answer that conflicts with the aggregate. If a model states a wrong price or capability, reproduce only the necessary claim and record the correct fact and source.

Prioritize Actions By Evidence And Business Value

Translate each finding into an action only when the team can name an owner and a success measure. A repeated factual error may justify correcting an owned page and contacting an inaccurate third-party source. A source gap may justify creating original evidence or a relevant expert contribution. A true product limitation belongs with product or sales enablement, not a content rewrite designed to obscure it.

Use a simple priority score: business importance from 1 to 3 multiplied by evidence strength from 1 to 3, divided by estimated effort from 1 to 3. The score helps order discussion; it is not a promise of outcome. Limit the weekly committed list to three actions so the report creates follow-through instead of an ever-growing backlog.

Carry unfinished actions forward with status and learning. If an action does not change the measured pattern after enough comparable observations, record that result. AI visibility work improves when failed hypotheses remain visible rather than being replaced by a new story each week.

  • Business importance: proximity to a valuable buyer decision and size of the affected segment.
  • Evidence strength: one isolated run, repeated prompt pattern, or repeated cross-platform pattern.
  • Effort: estimated work and dependencies, scored consistently.
  • Success measure: a component metric, corrected fact, relevant earned source, qualified visit, or downstream business outcome.

Worked Example: A Report That Does Not Overclaim

Assume the core set contains 30 prompts with three runs each in ChatGPT Search and Claude web search. ChatGPT produces 88 valid runs and Claude 90. Your brand appears in 26 ChatGPT runs and 18 Claude runs, so mention rates are 26/88 = 29.5% and 18/90 = 20.0%. In the prior comparable period they were 21/88 = 23.9% and 19/90 = 21.1%. Report the ChatGPT increase and the essentially flat Claude result separately.

The ChatGPT movement is concentrated in four integration prompts. The same two independent comparison pages appear as visible sources in nine of the improved observations. The correct finding is that visibility increased within the monitored ChatGPT sample for that cluster and the pages are a recurring source pattern. The incorrect finding is that those pages delivered a 5.6-point increase or that overall buyer awareness grew.

The action list might be: verify that both publisher pages describe the integration accurately; strengthen the owned integration page with current compatibility evidence; and rerun the unchanged cluster next week. If Search Console separately reports more Google generative impressions for the integration page, show that in the first-party panel as corroborating topic-level evidence, not prompt-level attribution.

Publish The Limitation Beside The Result

End with the methods statement every week, even when nothing changed. The report represents a controlled sample of generated answers, not all user prompts, platform traffic, market share, or guaranteed future outputs. Product behavior, retrieval, personalization, locale, and model changes can affect results.

Google's Search Console report has a different limitation: it is first-party within its supported features, but it does not reveal exact prompts, answer language, competitor mentions, or causal reasons for page selection. Keeping both limitations visible makes the report useful for decisions without pretending that an emerging channel is fully observable.

Copy-Ready Template

1. Report Header And Data Quality

  • Reporting period: [YYYY-MM-DD to YYYY-MM-DD] | Comparison: [previous week / four-week baseline] | Owner: [name] | Status: [final / directional / partial].
  • Scope: [platforms and modes] | Market: [country] | Language: [language] | Core prompt version: [version] | Competitor set: [version].
  • Runs expected: [#] | Valid: [#] | Invalid: [#] | Missing: [#] | Repetitions per prompt: [#].
  • Comparability note: [No material method changes / describe prompt, model, mode, locale, source-parsing, or access change and affected metrics].
  • Reviewer check: [name] reviewed [# or %] of observations, all material inaccuracies, and every finding used in the executive summary.

2. Executive Summary

  • Overall: Within the monitored [platform/mode] sample, [metric] changed from [numerator/denominator and %] to [numerator/denominator and %]. The movement was concentrated in [prompt cluster], while [other cluster/platform] was [flat/mixed/down].
  • Strongest evidence: Across [#] comparable observations, [specific repeated pattern]. Supporting observations: [IDs or internal link].
  • Material risk: [#] observations contained [outdated/inaccurate/unverifiable] claims about [fact]. Correct information and source: [fact + URL].
  • Source pattern: [domain/page type] appeared in [#] relevant observations, especially for [cluster]. This is an association to investigate, not proof of causation.
  • This week's decision: [continue monitoring / correct information / improve evidence / investigate product gap / no action] because [brief evidence-based reason].

3. Scorecard

  • Metric | Platform/segment | This period (count and rate) | Comparison (count and rate) | Change | Interpretation.
  • Unbranded mention rate | [platform + cluster] | [#/# = %] | [#/# = %] | [± points] | [directional finding].
  • Top-three rate | [qualifying recommendation prompts] | [#/# = %] | [#/# = %] | [± points] | [ordered lists only].
  • Share of voice | [frozen competitor set] | [#/# appearances = %] | [#/# = %] | [± points] | [largest brand movement].
  • Owned citation rate | [web-sourced mode] | [#/# opportunities = %] | [#/# = %] | [± points] | [leading owned page].
  • Material inaccuracy rate | [mentioning observations] | [#/# = %] | [#/# = %] | [± points] | [affected fact].
  • Search Console generative impressions | Google Search only | [#] | [#] | [% or absolute] | [top page/country/device; keep separate from sample rates].

4. Prompt-Cluster And Competitor Review

  • Cluster: [name] | Business weight: [1–3] | Prompts: [#] | Valid runs: [#] | Mention rate: [#/# = %] | Top-three rate: [#/# = %].
  • Buyer question represented: [the decision this cluster tests].
  • Recurring competitors and positions: [brand — # appearances — typical context].
  • What changed: [specific count-based observation compared with the same prompt version].
  • Raw evidence: [observation IDs or internal link].
  • Next check: [hypothesis to test without claiming causation].

5. Citation And Accuracy Log

  • Prompt ID | Platform | Run | Brand outcome | Source URL | Source type [owned/earned/directory/forum/other] | Claim supported | Reviewer note.
  • Accuracy issue ID | Observation ID | AI claim | Classification [outdated/inaccurate/unverifiable] | Correct fact | Primary source URL | Material impact | Owner.
  • Source gap: For [high-value prompt], [competitor] appeared and [our brand] did not. Recurring cited domains: [domains]. Evidence or qualification to investigate: [specific topic].
  • Concentration note: Top three relevant domains supplied [#/# = %] of relevant domain appearances, counting each canonical domain at most once per observation. Interpretation: [healthy diversity / concentration risk / insufficient sample].

6. Three-Action Plan

  • Action 1: [specific action] | Evidence: [finding/IDs] | Owner: [name] | Due: [date] | Success measure: [metric or verified outcome] | Priority score: [importance × evidence ÷ effort].
  • Action 2: [specific action] | Evidence: [finding/IDs] | Owner: [name] | Due: [date] | Success measure: [metric or verified outcome] | Priority score: [score].
  • Action 3: [specific action] | Evidence: [finding/IDs] | Owner: [name] | Due: [date] | Success measure: [metric or verified outcome] | Priority score: [score].
  • Carried forward: [action] | Status: [not started/in progress/blocked/complete] | Learning so far: [what the evidence now suggests].
  • Backlog, not committed this week: [item + reason].

7. Copy-Ready Methods And Limitations Note

  • This report summarizes a controlled sample of generated answers collected from [platforms/modes] in [country/language] between [dates]. The core set contained [#] prompts at version [version], run [#] times per platform. Rates use valid runs and show counts beside percentages. Branded diagnostic prompts are excluded from unbranded discovery metrics.
  • Results are directional observations, not all user prompts, platform traffic, market share, deterministic rankings, or guaranteed future answers. Generated outputs may vary by run, model, mode, retrieval, locale, personalization, and product changes. Material method changes are annotated in the data-quality section.
  • Google Search Console generative AI impressions, where included, are reported in a separate first-party panel for supported Google Search features. They are not combined with sampled ChatGPT or Claude answer rates, and topic-level patterns are not treated as one-to-one prompt attribution.

Automate the recurring work

Replace manual collection with a traceable weekly report

Keep the raw answers, calculations, competitor gaps, cited sources, and resulting publishing actions connected.

Explore automated AI visibility reports