Machine Relations

AI Visibility Scores Are Not Comparable: A 9-Platform Methodology Audit

A primary-source audit of nine AI visibility platforms found three vendors publishing contradictory metric definitions, three consistent contracts, and three sets of complementary metrics that buyers routinely confuse.

Jaxon Parrott
Jaxon ParrottSep 10, 2026

AI visibility scores from different tools are not comparable numbers.

A score of 50 can mean your brand appeared in half of all answers, half of answers that mentioned any brand, half of an impression-weighted competitive set, or a composite of topic coverage and mention consistency. Those are different measurements. Moving the same prompt set between platforms can change the score even when the underlying answers do not change.

We audited the public methodology, help, glossary, product, FAQ and API pages of nine AI visibility platforms on September 10, 2026. Three vendors publish internally contradictory definitions for a headline metric. Three publish definitions that agree across the surfaces we checked. Three publish multiple metrics that are defensible but easy to compare incorrectly.

The conclusion is not that every platform is unreliable. It is that the measurement contract matters more than the score.

The audit result

VendorPublic measurement contractVerdict
ProfoundVisibility is described as brand-containing responses divided by either all tracked responses or only responses containing at least one brandContradictory
Ahrefs Brand RadarAI Share of Voice is described as impression-weighted share in help documentation and as share of responses in product FAQ copyContradictory
SimilarwebProduct and KPI pages separate response visibility from mention share; one vendor blog FAQ defines mention share with a response denominatorContradictory
EvertuneVisibility is consistently the share of responses mentioning the brand; AI Brand Score adds position weightingConsistent
Peec AIVisibility is consistently brand-mentioning responses divided by total responses, including in its API exampleConsistent
DaydreamShare of answers naming a brand is computed per tool and averaged across five toolsConsistent
Scrunch AISeparate metrics intentionally use all responses, citing responses, citation URLs or competitive mentions as denominatorsComplementary
SemrushAI Visibility is a composite of topic coverage and mention consistency, while overview copy uses simpler presence languageComplementary
ProminaraPrompt-run visibility and a site-readiness score are separate instruments with separate unitsComplementary

“Contradictory” here means two public pages from the same vendor state measurement rules that can produce different numbers from the same dataset. “Complementary” means different denominators are attached to explicitly different metrics. “Consistent” means the public definitions we checked agree; it does not certify private implementation.

Prominara is operated inside the AuthorityTech portfolio. It is included to make our own measurement contract inspectable, not as independent validation of it.

Profound publishes two denominators for the same Visibility Score

Profound provides the cleanest demonstration of why a metric name is not a measurement contract.

Its Answer Engine Insights overview defines Visibility Score as:

responses that include your brand
÷ responses that include at least one brand

Its glossary gives a worked example based on 50 brand appearances out of 100 total tracked responses:

responses that include your brand
÷ all tracked responses

Both pages showed “Last updated 2 months ago” when checked. This is not an obvious current-versus-legacy documentation conflict.

The difference gets large precisely when visibility is weak. Suppose 100 answers are tracked, 50 mention your brand and only 60 mention any brand. The all-response denominator produces 50%. The brand-containing-response denominator produces 83.3%.

Nothing about the brand changed. The denominator removed 40 failures.

This does not establish which formula Profound’s product currently executes. It establishes that a buyer cannot determine the formula from Profound’s public definition alone because Profound publishes both.

Profound also states that prompts run once per day per configured model. That run count is part of the contract too: once-daily sampling measures a different distribution from repeated runs of every prompt.

Ahrefs describes AI Share of Voice as both impressions and responses

Ahrefs’ public material is precise about several lower-level events. Its AI Visibility Metrics documentation counts a brand mention once when the brand appears at least once in an answer, and a domain citation once when at least one page from that domain is cited. Those are clear deduplication rules.

The contradiction appears at the headline layer.

The same help page defines AI Share of Voice as a brand’s percentage share of impressions compared with tracked brands. Ahrefs’ Brand Radar methodology explains that estimated impressions weight mentions using Google search volume and warns that the result represents potential visibility rather than observed audience reach.

But the Brand Radar product FAQ describes AI Share of Voice as the percentage of AI responses in a topic set that mention or cite the brand versus competitors.

An impression-weighted share and an unweighted response share are not interchangeable. A brand concentrated on high-volume prompts can lead the first and trail the second. If the FAQ is shorthand, it does not say so.

Ahrefs also exposes a valuable distinction most tools do not: pages retrieved in the background but not cited can still be recorded as “Found in” an AI response. That separates retrieval from visible citation and makes the lower-level data more operationally useful than the ambiguous headline label.

Similarweb’s product contract is clear; one vendor FAQ overwrites it

Similarweb’s AI Brand Visibility product page defines two different metrics:

  • Brand Visibility: answers mentioning the brand divided by all tracked answers.
  • Brand Mention Share: the brand’s mentions divided by all brand mentions across tracked topics.

Its GEO KPI guide supplies worked examples that preserve the split: 1,346 brand-containing answers out of 5,556 total answers produces 24.23% visibility, while the same 1,346 mentions out of 47,020 total brand mentions produces 2.86% mention share.

Its AI Share of Voice guide makes the same distinction.

A separate Similarweb AI consumer journey FAQ defines Brand Mention Share as the percentage of relevant AI responses containing a brand name. That is the response denominator used for visibility, not the competitive-mention denominator used elsewhere for mention share.

The product and KPI pages form a coherent contract. The FAQ conflicts with it. A buyer copying the FAQ definition into an internal dashboard would rebuild the wrong metric while citing Similarweb accurately.

That is the dangerous class of documentation error: the sentence is real, the citation is real and the resulting implementation is still wrong.

Three platforms publish consistent public definitions

Consistency is possible, and three vendors demonstrate it across the pages checked.

Evertune: responses first, position second

Evertune’s metrics documentation, glossary and public FAQ all define Visibility Score as the percentage of AI responses mentioning the brand. AI Brand Score is a separate metric that weights visibility by answer position.

Evertune also states that it samples every prompt 100 times per model and claims that this produces roughly a one-point margin of error at the overall level and two points at topic level. That uncertainty claim is Evertune’s, not an independently reproduced result, but the disclosed run count is unusually specific.

Peec AI: the documentation and API arithmetic agree

Peec AI publishes the direct formula in its visibility documentation: responses mentioning the brand divided by total responses.

Its brands report API documentation uses the same aggregation rule, sum(visibility_count) / sum(visibility_total), and gives a worked example of 5 divided by 10 producing 0.5.

That agreement matters because prose and API arithmetic often drift. Here they do not.

Daydream: averaged across tools, not pooled

Daydream’s leaderboard methodology defines presence as a brand name appearing in an answer, then averages each tool’s answer share across five AI tools rather than pooling every answer into one denominator.

Its compliance automation category page declares the frame: 12 questions × 3 phrasings × 3 runs × 5 tools = 540 answers. The arithmetic checks, and the page repeats that results are averaged across tools.

Averaging and pooling can produce different numbers when tools return different numbers of valid answers. Daydream states which operation it uses.

Different denominators are not a defect when the metric names stay different

Scrunch AI is the strongest example of multiple denominators used deliberately rather than accidentally.

Its metrics guide defines mention rate and citation rate against all responses. Its citation metrics documentation separately defines:

  • citation URL share: your citation URLs divided by all citation URLs;
  • citation share of voice: responses citing your domain divided by responses citing any source;
  • citation rate: all responses citing your domain divided by all responses.

Its Data Studio documentation exposes the response-level aggregation, and its metrics reference records metric-specific denominator rules.

Those values should differ. The documentation tells the buyer why.

Semrush presents a different kind of complement. Its AI Visibility Data documentation says AI Visibility combines Topic Coverage and Mention Consistency into a normalized score. Its Visibility Overview describes the result more loosely as overall share of presence and how often the brand is mentioned. The simplified wording is incomplete, but it does not publish a second worked formula that contradicts the composite definition.

Prominara separates two instruments that share visibility language. Its methodology measures repeated unbranded prompt runs, records whether retrieval occurred and decomposes citation probability as Pr(retrieval) × Pr(cited | retrieval). Its public site check is a weighted readiness audit of page and site factors. One measures answer-engine outcomes; the other measures source eligibility. Comparing their scores would be a category error even though both are labelled around AI visibility.

The seven questions to ask before buying an AI visibility tool

A buyer does not need every internal implementation detail. A buyer does need enough information to reconstruct what a number means.

Ask these seven questions before comparing vendors or setting a baseline:

  1. What enters the denominator? All answer runs, only successful answers, only answers naming a brand, all brand mentions, all citations, or estimated impressions?
  2. What is the unit? Prompt, run, answer, mention, citation URL, citing response, domain or topic?
  3. How many runs are collected per prompt and engine? One daily draw and 100 repeated draws estimate different distributions.
  4. How are engines combined? Pooled across every answer, averaged equally by engine, or weighted by modeled impressions?
  5. What gets deduplicated? A brand mentioned three times in one answer can be one response event or three mention events.
  6. Are retrieval and citation separated? A page can be retrieved and rejected. A cited page is a subset of retrieved pages, not a synonym.
  7. What uncertainty is reported? At minimum, the numerator, denominator and sample size. Prefer intervals or repeated-run stability.

If a vendor cannot answer those questions, do not compare its score with another platform’s score. Treat it as a private index whose movement may be useful inside the product but whose level is not portable outside it.

Do not buy a score. Buy a measurement contract.

AI visibility tools observe a probabilistic system through different prompt sets, engines, locales, cadences and sampling intensities. Variation is unavoidable. Ambiguity is not.

A useful platform can use any defensible denominator if it states the unit, scope, aggregation and exclusions consistently. A misleading platform can use a mathematically valid formula and still make the result unreadable by changing the denominator between pages.

The buyer standard is simple:

metric name
+ numerator
+ denominator
+ unit of observation
+ sampling frame
+ aggregation rule
+ uncertainty

Without that contract, a score is not comparable evidence. It is a dashboard-specific label.

That distinction matters beyond procurement. If a leadership team changes platforms and its AI Visibility Score rises from 32 to 61, it may have improved. It may also have moved from an all-response denominator to a brand-containing-response denominator, from pooled engines to equal-weighted engines, or from response share to impression-weighted share.

The number moved. Until the contract is known, the brand may not have.

Method and limits

This audit checked public vendor-controlled pages on September 10, 2026: help centers, documentation, glossaries, methodology pages, FAQs, product pages and API references. We searched for every public definition and worked example we could verify for the named headline metrics rather than stopping at the first matching passage.

The audit does not inspect private application code or certify what any platform computes internally. Public documentation can lag a product, and some vendors may expose more detail inside authenticated interfaces. “Consistent” means the public pages checked agreed. “Contradictory” means public pages stated rules that can produce different values from the same observations. “Could not verify” was never treated as evidence that a capability or definition does not exist.

The claims above are limited to the nine named vendors and the cited pages as they appeared on the audit date. They do not establish the state of every AI visibility product or guarantee that the pages will remain unchanged.

FAQ

What is the best AI visibility tracking tool?

There is no universal best tool because the platforms measure different units. Choose against the decision you need to make. Use response-level presence for brand discovery, citing-response rate for source selection, citation URL share for source competition, and a retrieval-aware instrument when you need to diagnose why citation rate changed. Require the vendor to disclose the measurement contract before you compare scores.

Why do two AI visibility tools give different scores?

They may use different prompt sets, run counts, engines, locales, denominators, deduplication rules or aggregation methods. Even identical answers can produce different scores when one platform divides by all responses and another divides only by responses that mention a brand.

Is AI Share of Voice the same as AI Visibility?

Not necessarily. Visibility often means the share of responses containing your brand. Share of Voice can mean your share of competitive mentions, responses, citations or impression-weighted exposure. Read the formula rather than the label.

Should AI visibility be measured once or with repeated runs?

Repeated runs are more informative because AI answers vary. A single run is one observation, not the underlying probability. If a platform runs once per prompt, use longer windows and avoid treating small day-to-day moves as strategy signals.

What should an AI visibility report include?

At minimum: numerator, denominator, sample size, prompt-set version, engines, locale, date window, runs per prompt, deduplication rule and per-engine results. A percentage without its counts is not auditable.