54% of AI Answers Name No Brand at All. That Is Why Two Visibility Scores Disagree
We ran 24 buyer prompts across four answer engines and measured how often the answer names no tracked competitor. It was 54%. That single number is the gap between the two denominators AI visibility platforms publish for the same metric.
Across 96 answers from four AI engines, 54% named no company from the tracked competitive set. That is the whole story of why two AI visibility platforms can report different scores for the same brand on the same prompts, and it is a quantity we did not find reported on any of the public methodology, help, glossary or API pages we audited across nine platforms.
Earlier today we audited how nine AI visibility platforms define their headline metrics and found that several publish internally contradictory definitions. The clearest case is Profound, which defines its Visibility Score two different ways on two of its own live help pages. Its glossary uses all tracked responses as the denominator: "McDonald's is mentioned in 50 responses out of a total 100 responses tracked. The visibility score will be 50% (50/100)." Its Answer Engine Insights overview uses a different one: "the number of responses that include your brand" divided by "the total number of responses that include at least one brand."
Documenting that contradiction is easy. The question a buyer actually has is: how much does it matter? Those two denominators differ by exactly the responses that mention nobody. So we measured how many of those there are.
The measurement
On 10 September 2026 we sent 24 prompts to four answer engines for 96 answers: Perplexity via its Sonar API, ChatGPT via the OpenAI API with web search, Gemini via grounding with Google Search, and Claude via the Anthropic web search tool. Every one returned usable text; there were no failures to exclude. These are retrieval-augmented API surfaces, not the consumer chat products, and they are the same class of surface commercial trackers query.
For each answer we checked the body text against a pre-registered set of 32 vendors covering the AI visibility and answer-engine-optimization category — among them Profound, Ahrefs Brand Radar, Semrush, Similarweb, Peec AI, Evertune, Otterly.AI, Scrunch AI, Conductor, and BrightEdge. An answer naming none of them is brand-free: it exists in the platform's tracked universe but names no member of the competitive set. This is the same share of voice construction the category borrowed from advertising, where the denominator has always been the contested part.
The prompt set was declared in three intent bands before the run, because the composition of a prompt set is not a detail here — it is the variable that determines the answer.
- Band A, vendor-seeking (8 prompts). "best ai visibility tracking tools 2026", "llm brand monitoring tool pricing". A named vendor is the expected answer.
- Band B, problem-aware (8 prompts). "how do i measure my brand visibility in ai search", "how to report ai search performance to executives". A vendor may or may not surface.
- Band C, conceptual (8 prompts). "what is answer engine optimization", "why does my brand not appear in ai answers". A vendor is not the natural answer.
The result
| Prompt band | Responses | Brand-free | Brand-free rate | Denominator divergence |
|---|---|---|---|---|
| A — vendor-seeking | 32 | 3 | 9% | 1.10× |
| B — problem-aware | 32 | 20 | 63% | 2.67× |
| C — conceptual | 32 | 29 | 91% | 10.67× |
| All 24 prompts | 96 | 52 | 54% | 2.18× |
The divergence column is what the two definitions do to the same underlying data. If brand mentions are held constant, dividing by branded responses instead of by all responses multiplies the score by 1 / (1 − brand-free rate). On this prompt set that factor is 2.18. A brand scoring 20% under one published definition scores 44% under the other, with not one answer having changed.
Per engine, the same measurement:
| Engine | Responses | Brand-free rate | Divergence |
|---|---|---|---|
| Gemini | 24 | 29% | 1.41× |
| Claude | 24 | 54% | 2.18× |
| Perplexity | 24 | 63% | 2.67× |
| ChatGPT | 24 | 71% | 3.43× |
Engine mix moves the number almost as much as prompt mix does. A platform weighting ChatGPT heavily and a platform weighting Gemini heavily will disagree on the same brand, the same prompts, and the same definition — before any methodology dispute begins.
What this actually means
The score is a property of your prompt set, not of your brand. This is the finding. A vendor-seeking prompt set produces a brand-free rate near zero, where the two denominators nearly agree and the score is close to honest under either reading. A conceptual prompt set produces a brand-free rate above 90%, where one denominator throws away nine answers in ten and inflates the score by more than a factor of ten. Most real tracked prompt sets are a mix, which is why most real scores sit somewhere in an unmarked band between the two readings.
"Improving your visibility score" and "appearing in more answers" are different projects. Under the branded-responses denominator, removing conceptual prompts from your tracked set raises your score without changing a single answer. That is not gaming; it is the arithmetic working as documented. But it means a rising score is not evidence of rising presence unless the prompt set held still, and nothing in the reported number tells you whether it did.
Comparing two vendors' scores is not meaningful without both denominators and both prompt sets. Two platforms reporting 31% and 68% for the same brand can both be correct. This is the concrete form of the point our methodology audit made in the abstract: the measurement contract matters more than the score.
Brand-free does not mean invisible. These answers were substantive; they explained concepts, gave procedures, and cited sources. They simply named no vendor. For a category still educating its market, the conceptual band is where the demand is — and it is precisely the band that one denominator discards. A platform can show your visibility improving while the answers where buyers form their understanding continue to name nobody.
Limits, and which way they bend
The tracked set is finite. An answer naming a vendor outside our 32 scores as brand-free, so the measured brand-free rate — and every divergence figure derived from it — is an upper bound. Two effects push the other way and make it conservative. ChatGPT inlines cited domains directly in its answer text, so a domain like conductor.com appearing as a citation marker counts as a brand mention in our method even where the prose names no vendor; that lowers its measured brand-free rate. And any false positive from an ambiguous vendor name does the same. Only seven of 96 answers named exactly one brand, so the result does not hinge on edge cases: the 44 branded answers named a median of 5 vendors each (mean 4.6, range 1 to 9), and the four most-named were Profound (32), Semrush (32), Otterly.AI (29) and Peec AI (27). Answers in this category either name a slate of vendors or name none, which is why the brand-free rate is a clean split rather than a gradient.
This is a single run on a single day. Answer engines are not deterministic — a spot re-query of one prompt produced a different vendor set on the second attempt. Per-cell figures therefore carry sampling noise, and the band-level pattern, which is monotone and large, is the robust part. Twenty-four prompts is a small set; the finding is the mechanism and its magnitude class, not a precise constant.
We measured English-language prompts in one category. Brand-free rates in categories with entrenched incumbents will be lower, and in emerging categories higher. Every buyer should run this on their own prompt set rather than adopt our 54%, which is why the method is three steps: send your tracked prompts, check each answer against your tracked competitor list, divide.
What to ask your vendor
One question settles it: which denominator does my score use, and what is my prompt set's brand-free rate?
A platform that can answer both can tell you what your number means. A platform that publishes both denominators on different pages cannot, and a platform that will not report your brand-free rate is reporting a ratio while withholding half of it.
The category has spent two years arguing about which visibility score is right. The measurement above suggests the argument is misdirected. The scores are not competing estimates of one quantity. They are different quantities, and the distance between them is a number you can measure in an afternoon.
Method: 24 prompts × 4 engines, 10 September 2026, 96/96 responses usable, 32-vendor tracked set, body text only. The measurement script is published with the underlying event record so the figures can be reproduced or contradicted.