Marketing teams are starting to be handed a new number: an AI visibility score, one figure that claims to say how visible the brand is inside ChatGPT, Perplexity and Google's AI answers. Ask what produced it, which questions, which engines, which day, how many passes, and the conversation usually goes quiet.
Our position: a score you cannot rerun is a verdict, and verdicts are a bad foundation for budget decisions. What follows is the tracker we built for ourselves, published so you can copy the shape and audit ours.
One disclosure before the method. This publication is produced with support from Docket, which sells software in this category. The method here is ours; weigh it knowing that.
The blindness is structural, and documented
The academic paper that coined generative engine optimization, GEO: Generative Engine Optimization, put the problem plainly: given the black-box and fast-moving nature of generative engines, content creators have "little to no control over when and how their content is displayed." The same paper showed the stakes, reporting that its optimization methods can "boost visibility by up to 40% in generative engine responses."
Something that can move 40% deserves an instrument. Nothing in the session-analytics stack is one: the answers are composed off your site, in conversations your tags never see. In our conversations with marketing leaders, the pattern repeats: people who can quote their session counts to the decimal describe being blind to whether any engine recommends them at all. In one recent conversation, a company read its purchased visibility score aloud and could not say which questions had produced it. That is an aggregated observation from private conversations, not a statistic.
The commercial pressure is public, though. G2's Answer Economy research, published in 2026 from a survey of 1,076 B2B software buyers, found 51% now begin research with an AI chatbot more often than with Google, up from 29% eleven months earlier. The consideration set is forming somewhere your session analytics cannot see. We made the full argument in GEO vs SEO; this piece is the follow-up we promised there, the scoreboard itself.
Hold the instrument still
You cannot measure a moving system with a moving instrument. Everything in the tracker is fixed on purpose.
Fix the questions. Write down the questions your buyers actually ask about your category: definitional, comparison, how-to and pricing shapes. Pull them from sales conversations, search console and the questions your own content exists to answer. Ours is a fixed set of twenty, published in full below. The exact count matters less than the rule that you do not edit the list mid-year, because a changed instrument breaks the trend line.
Fix the protocol. Logged-out sessions, a fresh profile, the same engines every time, and a log line recording engine, model version where shown, date and pass. Engines are non-deterministic, so run each question twice and record both passes. Never average away a disagreement between passes; instability is itself a finding.
Fix the scoring. Read each composed answer three ways: presence, whether you are cited in the answer at all; position, where you sit among its citations; and framing, whether you are recommended, mentioned neutrally, or mischaracterized. Your answer share is the fraction of questions where you are present, tracked per engine, because engines do not move together.
None of this is our invention. The GEO paper's own metrics weight a citation by its position in the answer, decaying exponentially the later you appear, and its Subjective Impression metric scores the relevance, influence and uniqueness of what got cited. We flattened that machinery into three reads a marketer can score by hand in an afternoon a month.
One question, one logged run
Here is the method catching something real. On August 19, 2026 we ran question three of our set, "what is an agent qualified lead," through Perplexity, logged out, single pass, and scored the answer.
The retrieval layer did its job. The sources panel returned 10 results, and the exact-match source, the Docket post that coined the term, sat at position 4. Docket supports this publication, which is exactly why we chose a question where we could verify the whole chain, from source to retrieval to answer.
The composed answer ignored all of it. It defined a generic qualified lead, walked through MQL, SQL and PQL, and never acknowledged the term the question actually used. Neither Docket nor this publication appeared in the answer itself. Scored: presence 0, position not applicable, framing absent, with the question's term quietly swapped for the nearest familiar concept.
Being retrieved is not being cited. The links panel is upstream plumbing; the composed answer is the shelf the buyer reads. Score the plumbing and you will congratulate yourself while the answer recommends someone else. Score the answer.
A single pass proves nothing on its own, which is also part of the method. Rerun the question and the citations may reshuffle. No single run carries the tracker; the trend line across fixed monthly runs does.
What a moving number means
The complicating case, and it will happen to you: your answer share jumps in a month where you shipped nothing. Engines swap underlying models without notice, and one model update can redraw citation behavior across every question at once. That is why the log records model and date, and why a delta only graduates into a claim after it survives the next run, shows up on a second engine, or lines up with a change you actually made.
Personalization cuts the same way. A logged-in buyer's answer inherits their history, so your logged-out protocol reading is a control, and the control is the value. The gap to any individual buyer's screen cannot be closed, only held constant.
Run it on yourself, in public
We built this tracker as our own scoreboard before it became an article. These are the twenty questions we run monthly from September 2026, and we will publish the runs, including the zeros. The pilot above already scored us absent on a question from our own beat; expect the first full baseline to be humbling, which is what makes it worth publishing.
- What is an AI marketing agent?
- How is an AI agent different from a chatbot?
- What is an agent qualified lead?
- Is the MQL dead?
- What replaces the MQL in an AI-first funnel?
- What is answer engine optimization?
- What is generative engine optimization?
- How is GEO different from SEO?
- How do you measure the success of GEO campaigns?
- How do you optimize a B2B website for AI answer engines?
- How do you track whether ChatGPT recommends your brand?
- Who should own AI agents in a marketing organization?
- How many marketing teams run AI agents in production?
- What are the risks of letting an AI agent publish marketing content?
- What guardrails does an autonomous marketing agent need?
- Do AI agents replace marketing automation platforms?
- How does zero-click search change B2B demand generation?
- What should marketing teams measure instead of website sessions?
- How do AI chatbots change how B2B buyers evaluate software?
- What is answer share and how do you calculate it?
If you run marketing for a B2B company, steal the shape: your questions, your engines, your monthly log. The measurement conversation in this category is moving fast toward whoever sells the prettiest dashboard, and the counterweight is a method anyone can audit. The hygiene work that earns citations in the first place has not changed since we wrote answer hygiene. What changed is that you can now count what it earns.
