What You'll Measure
Give your team a shared vocabulary for four measurement pillars before touching a single report — strategy without a baseline is guesswork.
Each pillar answers a different question, uses a different report, and — this is the part worth emphasizing — none of them substitute for each other. A brand can score well on one and poorly on another, and the gap between them is often the most useful finding of the whole program. Each pillar below is marked with which starting-point tier it's available in.
Pillar 1: External AI Citations — Citation Monitor
Available: from day one
Question it answers: When a prospect asks ChatGPT, Claude, or Gemini directly, does your brand get recommended — or does a competitor?
Citation Monitor runs tracked queries through five signals — ChatGPT and Claude (closed-book training data), Google Gemini and ChatGPT Web Search (web-grounded), and optional Google organic rank — then scores a Citation Health Score, a verdict tier, and a diagnostic quadrant per query.
What to show your team: the Competitor Share of Voice table and the Google vs. Bing Citation Gap breakdown. Both make an abstract "we're not visible" complaint into something specific and actionable.
Real example to reference: a real run scored 32/100 ("Emerging"), with the brand cited only on its own branded query while a competitor (Ada) held 19% share of voice and won all three uncited category queries outright. The report's own diagnostic didn't stop at "you're not cited" — it named the exact competitor beating them and the exact pages to build in response.
Pillar 2: Real Customer Questions — Q&A Analysis
Available once you have real chatbot conversation history
Question it answers: What are people actually asking your chatbot, and how well is it answering?
Q&A Analysis clusters real chatbot conversations into up to 50 themes, computes an IDK (unanswered question) rate, and scores answer quality per theme. It's the only pillar built from your own real conversation data rather than a simulated audit.
What to show your team: the Query Frequency Histogram and Citation Phrase Suggestions panel — the mechanism that turns "what our customers actually type" directly into a Citation Monitor tracked-query list with one click. Between full runs, Q&A Insights shows the same signal live, per conversation.
Real example to reference: a Q&A Analysis run that found 30 out of 42 total consolidated action items traced back to Q&A Analysis findings alone — including five separate FAQ answers ranked High-impact/Low-effort because they were high-volume, unanswered, or answered inconsistently.
Pillar 3: Keyword Answer Quality — Keyword Visibility
Available once your bot is live — no history required
Question it answers: When someone asks a short, direct question vs. a longer conversational one, does your answer quality hold up either way?
Keyword Visibility expands seed keywords into natural-language questions across buyer-intent stages, scores every answer against your live chatbot, and computes a Format Delta between short-form and conversational query performance.
Don't treat this as a one-time score — it's the tool for proving a CMS content edit actually helped: re-run the same Keyword List after your team updates a page, and compare. Source the seed keywords from something real when possible — Google Search Console/Analytics query data, your old on-site search logs, or Q&A Analysis's Citation Phrase Suggestions — rather than guessing.
What to show your team: a real Format Delta example is the single best "aha" moment in this pillar — most people assume short, direct questions perform best, and the data frequently says the opposite.
Real example to reference: a real audit scored a short-form average of 82 against a conversational average of 89 — a +7.0 delta, meaning longer, context-rich questions actually scored higher, because the bot could synthesize multiple facets into one coherent narrative instead of hedging on a single direct yes/no.
Pillar 4: AI Referral Traffic — AI Referral Analytics
Available once you have real AI-referred traffic to attribute
Question it answers: Is any of this actually driving real visits — and once someone arrives, what do they do?
AI Referral Analytics is the one pillar that isn't a generated report — it's a live dashboard tracking sessions by engine, landing page, and engagement, with a drill-down into the actual chatbot conversation that followed. This is also where Perplexity shows up — it isn't one of Citation Monitor's five direct signals, but it's fully tracked here as a referral source.
What to show your team: the Sessions by Answer Engine donut and the Latest AI Referral Events feed — the fastest way to make "AI referral traffic" feel like a real, countable thing rather than a marketing phrase.
Real example to reference: a real dashboard view showed 7,842 AI referral sessions (up 28.6% period over period), with ChatGPT alone responsible for 43.5% of them — a single, concrete number that reframes "should we care about GEO" as "look how much of our traffic already comes from it."
The Pattern to Name Explicitly
Once relevant pillars are on the table, say this out loud: these pillars measure different failure modes, and a brand can fail at one while succeeding at another. A brand can have excellent on-site answers (strong Q&A Analysis or Keyword Visibility scores) while scoring "Emerging" on Citation Monitor — meaning the content is good, but nobody outside the website has found it yet. That specific pattern — strong on-site answers, weak external citation — is one of the most common findings this program surfaces, and it points squarely at content amplification and technical discoverability work, not a content-quality problem at all.
Related Documentation
- Where You're Starting From — Full breakdown of what's available at each stage
- Citation Monitor
- Q&A Analysis
- Keyword Visibility
- AI Referral Analytics