Measurement framework
Why Most AI Visibility Reports Are Wrong: The Measurement Problem in GEO
A data-driven framework defining the core metrics for generative visibility.

A screenshot is not a measurement system
One prompt, one model, and one screenshot can be useful as a qualitative observation. It cannot support a reliable trend, competitor claim, or ROI decision. Generative answers vary by prompt wording, model, retrieval context, account, geography, date, and random sampling.
A defensible measurement framework
Let a study define a prompt set P, models M, competitors C, and observation dates T.
| KPI | Definition | Formula |
|---|---|---|
| Prompt Coverage | Share of intended prompts successfully run | valid prompt runs / planned runs |
| Brand Mention Rate | Share of valid answers naming the brand | answers mentioning brand / valid answers |
| Citation Rate | Share of brand mentions supported by a source | cited mentions / brand mentions |
| Recommendation Rate | Share of answers recommending the brand or its service | recommendations / relevant valid answers |
| Share of AI Voice | Brand mentions divided by all tracked brand mentions | brand mentions / all tracked mentions |
| Competitor Visibility | Same visibility measure for each competitor | competitor mentions / valid answers |
| Model Variance | Spread of a KPI across models | max KPI − min KPI, or standard deviation |
| Temporal Variance | Spread across collection dates | max KPI − min KPI across T |
The denominator must be published. A failed run, an answer that does not address the category, and an answer with no candidate list should not be silently counted as zero.
Design the dataset before the dashboard
Build prompts from user intent: definition, comparison, recommendation, local need, problem diagnosis, and purchase investigation. Keep a stable core set and a rotating exploratory set. Capture exact outputs, citations, model metadata, locale, and timestamps. Hashing or versioning the prompt set prevents accidental methodological drift.
Interpretation guardrails
Mention is not authority. Citation is not endorsement. Recommendation is not conversion. A high score in one model is not market dominance. Report confidence intervals when the sample permits, and show raw counts beside percentages.
From report to product
A future AI Visibility dashboard should let a team filter by intent, language, model, date, competitor, source domain, and citation quality. It should highlight changes, not manufacture a single magic score. The correct output is a decision queue: which evidence is missing, which prompt family is weak, and what should be tested next.
Read the strategic context in GEO Is Not the New SEO and the authority system in Evidence Architecture.