Layer 1: visibility in AI answers

Use a fixed set of buyer-relevant prompts and repeat it across relevant engines, markets, and dates. Track prompt mention coverage (prompts where the brand appears divided by prompts tested), response mention rate (responses containing the brand divided by responses recorded), and owned-domain citation rate (responses citing your domain divided by responses recorded).

Also track mention share of voice (your mention events divided by all tracked competitor mention events), accuracy rate (accurate brand descriptions divided by reviewed brand descriptions), and cited-page breadth (the number of distinct owned pages cited). Document whether the denominator is prompts, responses, citations, or estimated impressions, since two vendors can use the same label for different calculations.

Layer 2: engagement and demand

AI assistants do not always send a click. Track direct AI referrals where referrer data is available, but also watch branded search, direct visits, return visits, assisted conversions, demo-page engagement, and self-reported discovery.

Add a “how did you hear about us” field with an AI assistant option, then preserve the free-text answer, and use campaign parameters for links you control. Do not claim that every rise in direct traffic came from AI. Google reports AI Overview and AI Mode traffic within the Web search type in Search Console rather than as a separate AI report, so combine Search Console trends with analytics and conversion data while acknowledging that attribution is incomplete.

Sources: Google Search Central: AI features

Layer 3: commercial outcomes

Connect sessions and self-reported discovery to qualified leads, pipeline, closed revenue, retention, or another outcome appropriate to the business, and compare conversion rate and sales quality by source where sample sizes allow.

A simple program-level return calculation is incremental gross profit attributed to the program, minus program cost, divided by program cost. The difficult part is attribution: use conservative rules and show them, and for small samples report counts and ranges instead of precise percentages that imply certainty.

Build a baseline that survives scrutiny

Freeze a core prompt set before reviewing the results, including problem, category, comparison, use-case, and brand prompts, and define competitors in advance. Store full responses and citations, not just extracted scores.

Run a quiet baseline over several cycles to observe normal variation. Compare changes only after the relevant work is live and validated, and segment by engine, since visibility can differ substantially between systems.

Connect work to outcomes

Maintain a change log with URL, finding, approved change, deployment date, validation result, target prompts or queries, and expected measure. This prevents teams from crediting a result to work that never shipped.

Rankout connects opportunities to approved work, validation, and results, which makes it possible to ask which completed actions preceded a visibility or conversion change instead of comparing two disconnected dashboards.

A useful reporting view

For leadership, show the core prompt coverage and citation trend, top accurate and inaccurate descriptions, competitors gaining visibility, cited-source gaps, AI-referred engagement, qualified outcomes, work completed, and next actions.

For practitioners, include raw responses, page-level citations, technical checks, prompt segments, implementation status, and measurement notes. Never hide methodology behind a single score.

Common measurement mistakes

Treating one favorable answer as persistent visibility, changing prompts every week and comparing incompatible samples, combining mentions and citations, and calling citations “rankings” are common errors that undermine a program's credibility.

Ignoring inaccurate or negative descriptions, reporting traffic without conversions or revenue without an attribution method, and measuring work before confirming it went live are equally common, and equally avoidable with a documented process.

What to take away

  • Track visibility, engagement, and commercial outcomes as three separate layers, not one blended score.
  • Freeze a core prompt set and competitor list before measuring, so results stay comparable over time.
  • Maintain a change log connecting approved work to deployment, validation, and the resulting metric.

Frequently asked questions

What is a good AI visibility score?

There is no universal benchmark. Scores depend on prompts, engines, markets, competitors, sampling, and formulas. Compare a documented baseline with itself over time.

Is a mention the same as a citation?

No. A mention names the brand; a citation identifies or links a source. Track both and review whether the surrounding description is accurate.

Can Google Search Console show AI Overview traffic?

Google includes traffic from its AI features within the Web search type. It does not provide a complete cross-platform view of ChatGPT, Claude, Gemini, and Perplexity.

How should an agency report AI visibility?

Show the prompt set, engines, dates, raw evidence, formulas, changes completed, and business outcomes. Keep proprietary scores secondary to reproducible measures.