Skip to main content

What We Measure in the Lab

The Lab teaches the measurement system first — because strategy without a baseline is guesswork.

Before touching a single fix, attendees need a shared vocabulary for four measurement pillars. Each one answers a different question, uses a different report, and — this is the part worth emphasizing — none of them substitute for each other. A brand can score well on one and poorly on another, and the gap between them is often the most useful finding of the whole session.


Pillar 1: Real Customer Questions — Q&A Analysis

Question it answers: What are people actually asking our chatbot, and how well are we answering?

Q&A Analysis clusters real chatbot conversations into up to 50 themes, computes an IDK (unanswered question) rate, and scores answer quality per theme. It's the only pillar built from a brand's own real conversation data rather than a simulated audit — which is why it's the natural starting point.

What to show the room: the Query Frequency Histogram and Citation Phrase Suggestions panel — the mechanism that turns "what our customers actually type" directly into a Citation Monitor tracked-query list with one click. This is the clearest illustration of the whole Lab's thesis: measurement pillars aren't siloed, they feed each other.

Real example to reference: a Q&A Analysis run that found 30 out of 42 total consolidated action items traced back to Q&A Analysis findings alone — including five separate FAQ answers ranked High-impact/Low-effort ("What does ai12z do," "What are the pricing plans," "How do I get started," "How do I add the bot to my website," "How secure is the platform") because they were high-volume, unanswered, or answered inconsistently.

Between full runs, Q&A Insights is the live, per-conversation table underneath this pillar — the same 1–100 quality score, visible per answer instead of per theme, with a one-click drill-down into the full transcript. See The Live Signal Reports for a dedicated walkthrough.


Pillar 2: External AI Citations — Citation Monitor

Question it answers: When a prospect asks ChatGPT or Gemini directly, do we get recommended — or does a competitor?

Citation Monitor runs tracked queries through five signals — ChatGPT and Claude (closed-book training data), Google Gemini and ChatGPT Web Search (web-grounded), and optional Google organic rank — then scores a Citation Health Score, a verdict tier, and a diagnostic quadrant per query.

What to show the room: the Competitor Share of Voice table and the Google vs. Bing Citation Gap breakdown. Both make an abstract "we're not visible" complaint into something specific and actionable.

Real example to reference: a real run scored 32/100 ("Emerging"), with the brand cited only on its own branded query while a competitor (Ada) held 19% share of voice and won all three uncited category queries outright. The report's own diagnostic didn't stop at "you're not cited" — it named the exact competitor beating them and the exact pages to build in response.


Pillar 3: Keyword Answer Quality — Keyword Visibility

Question it answers: When someone asks a short, direct question vs. a longer conversational one, does our answer quality hold up either way?

Keyword Visibility expands seed keywords into natural-language questions across buyer-intent stages, scores every answer, and computes a Format Delta between short-form and conversational query performance.

What to show the room: a real Format Delta example is the single best "aha" moment in this pillar — most teams assume short, direct questions perform best, and the data frequently says the opposite.

Real example to reference: a real audit scored a short-form average of 82 against a conversational average of 89 — a +7.0 delta, meaning longer, context-rich questions actually scored higher, because the bot could synthesize multiple facets (pricing tiers, features, usage) into one coherent narrative instead of hedging on a single direct yes/no.


Pillar 4: AI Referral Traffic — AI Referral Analytics

Question it answers: Is any of this actually driving real visits — and once someone arrives, what do they do?

AI Referral Analytics is the one pillar that isn't a generated report — it's a live dashboard tracking sessions by engine, landing page, and engagement, with a drill-down into the actual chatbot conversation that followed.

What to show the room: the Sessions by Answer Engine donut and the Latest AI Referral Events feed — the fastest way to make "AI referral traffic" feel like a real, countable thing rather than a marketing phrase.

Real example to reference: a real dashboard view showed 7,842 AI referral sessions (up 28.6% period over period), with ChatGPT alone responsible for 43.5% of them — a single, concrete number that reframes "should we care about GEO" as "look how much of our traffic already comes from it." For the individual-visitor version of this same signal, Q&A Insights shows the actual conversation a referred visitor had, tagged with the exact engine, landing page, and device metadata detected for that session.


The Pattern to Name Explicitly

Once all four pillars are on the table, say this out loud: these four pillars measure four different failure modes, and a brand can fail at one while succeeding at another. A brand can have excellent Q&A Analysis scores (customers get good answers on-site) while scoring "Emerging" on Citation Monitor (external AI doesn't know or cite them) — meaning the content is good, but nobody outside the website has found it yet. That specific pattern — strong on-site answers, weak external citation — is one of the most common findings the Lab exercise surfaces, and it points squarely at content amplification and technical discoverability work, not a content-quality problem at all.