The AI visibility tool market formed fast, and the pitches are confident. Profound leads its homepage with "Optimize Your Brand's Visibility in AI Search," over the claim that more than 100 million people search with AI every day. AthenaHQ opens with "Become the Brand AI Trusts" and calls itself "the command center for AEO and GEO." Both advertise coverage across the full engine roster: ChatGPT, Perplexity, Gemini, Claude, Copilot, Grok and Google's AI surfaces.
One disclosure before anything else. This publication is produced with support from Docket, which sells software in this category. That is exactly why this piece scores dimensions and never ranks vendors: what follows are the five questions we would put to any tool in this market, and how to check the answers yourself.
Either vendor above may be excellent. We have not tested them, and nothing here scores them. The questions exist because the difference between a measurement instrument and a rented scoreboard is structural, and a demo is built to hide structure.
Why a tool at all
The blindness these products address is real and documented. The academic paper that coined generative engine optimization, GEO: Generative Engine Optimization, put it plainly: given the black-box and fast-moving nature of generative engines, content creators have "little to no control over when and how their content is displayed," while its optimization methods moved visibility "by up to 40%." Something that can move 40 percent deserves an instrument, and nobody wants to score answers by hand forever. We built a manual answer-share tracker for exactly this, and the honest case for buying a tool is automating that kind of method without changing what it measures. The five questions test whether a product does the first thing or quietly abandons the second.
One: whose questions does it run?
A visibility score is computed over some set of prompts. Ask whether the tool runs your fixed question set, imported and locked, or its own canned panel. If it cannot take your questions, it is measuring its market, not your buyers. And a panel you cannot see or edit cannot be held still, so it cannot produce a trend line you can defend.
Two: does it score the answer or the plumbing?
Engines retrieve sources, then compose an answer, and the two layers disagree. In our own logged pilot run on August 19, 2026, the exact-match source for "what is an agent qualified lead" sat at position 4 in Perplexity's sources panel while the composed answer ignored the term entirely and cited nobody involved. A tool counting retrievals would have scored that a win. The buyer never reads the sources panel; the buyer reads the answer. Ask which layer the presence number counts, and make the demo show one raw example end to end.
Three: which engines, and which model versions?
Engine coverage lists are the easy part, and every vendor advertises a long one. The harder question is whether each logged run records the model version and the date. Engines swap underlying models without notice, and one swap can redraw citation behavior across your entire question set overnight. A number without a model log cannot tell you whether you changed or the engine did.
Four: can you re-score it?
An auditable score travels with the answer text that produced it. The GEO paper's own metrics weight a citation by its position in the answer and score the character of what got cited, which is only possible while the raw answer is in hand. If a tool exports the number but not the stored answers, you cannot re-read presence, position or framing, and you are back to trusting a verdict. Exportable raw runs are the difference between an instrument and an oracle.
Five: what does it do with a jump?
Your answer share will move in months where you shipped nothing. The tool's job is to stop you from writing that as your win or your crisis: per-engine trend lines because engines do not move together, repeated passes so instability is visible instead of averaged away, and annotations when an engine swaps models. If the demo shows one blended score climbing smoothly, ask what it is averaging.
Bring your zeros to the demo
The cheapest diligence in this market costs an afternoon. Run your own twenty questions by hand once, logged out, and score them before you see any product. Then ask the vendor to run the same questions and reproduce your zeros. Where their number disagrees with your logged run, one of you is measuring something else, and the demo's job is to find out which. A vendor comfortable with that conversation is selling an instrument.
None of this replaces the work the scoreboard watches. The citations these tools count are earned by answer hygiene that has not changed since we wrote it down: answer directly, source your claims, structure for extraction. Buy the instrument when your method outgrows your afternoon. Just make sure it is your method the tool is automating.
