The AI Share of Voice Measurement Contract: Prompt Set, Providers, Aliases, and Denominator
The eight decisions to fix before you collect a single AI answer — prompt set, provider, model, response mode, sampling window, entity aliases, denominator, and uncertainty rules — and how to write them down so two measurement cycles are comparable.
A share-of-voice number is the output of a measurement contract, so the contract has to be written before the first answer is collected. Eight decisions set it: what prompts were tested, which providers and product modes answered them, which model and response mode, over what sampling window, which brand aliases counted, how repeated runs were handled, what denominator turned observations into a percentage, and how uncertainty is reported. Change any one of them and the same brand returns a different number from the same underlying answers.
For the definition, the formula, and why one blended figure is the wrong unit, start at AI Share of Voice: Definition, Formula, and How to Measure It by Question Shape. This page is the contract that comes next: it fixes each of the eight decisions, shows a reproducible audit, and separates mention breadth from citations, links, recommendations, referrals, conversion, pipeline, and revenue. The companion reporting protocol covers sampling, auditable counting rules, and the per-cycle report. Machine Relations uses AI SOV as one measurement layer inside a broader system for building and testing machine-visible authority.
Define the measurement contract first
Before calculating AI SOV, write the contract that another analyst could repeat. A complete AI SOV record includes these fields:
| Field | Decision to record | Why it matters |
|---|---|---|
| Metric | Mention share, prompt-level mention incidence, rank position, share of citations, link inclusion, referral traffic, conversion, pipeline, or revenue | Each metric answers a different question and uses a different denominator. |
| Numerator | The count assigned to your brand: brand mentions, prompts with at least one mention, cited source URLs, linked URLs, sessions, opportunities, or revenue | One response can contain several brand mentions, one mention, zero citations, or several citations. |
| Denominator | Total brand mentions, total prompts, total citation slots, total linked URLs, total sessions, or total opportunities in the same sampled universe | Changing the denominator changes the metric even when the underlying responses are identical. |
| Sampling unit | Response, prompt run, prompt-provider pair, prompt cluster, account, locale, or day | The same prompt can vary across runs, accounts, locales, and product modes. |
| Prompt universe | The exact prompt library, prompt category, inclusion rule, excluded prompts, and version of the library | Category prompts, comparison prompts, best-of prompts, support prompts, and buyer-problem prompts do not measure the same opportunity. |
| Provider and model | Provider, model or product name, version if visible, browsing or retrieval mode, and citation display mode | ChatGPT, Perplexity, Claude, Gemini, Google AI Overviews, Copilot, and Grok should not be blended unless the contract says how they were weighted. |
| Locale and account state | Country, language, device if relevant, signed-in state, personalization state, and any enterprise workspace constraints | Personalization, geography, and enterprise context can change outputs. |
| Repetition policy | Number of runs per prompt, time between runs, randomization order, and retry rules for refusals or tool errors | A single run is an observation, not a stable estimate. |
| Observation window | Collection date, time zone, start/end dates, and whether the report is a point-in-time baseline or a rolling window | AI products and source indexes change; old readings should not be described as current platform behavior. |
| Entity-alias rule | Accepted names, excluded homonyms, product-vs-company mapping, merger names, and manual adjudication rule | "AuthorityTech," "Authority Tech," and a product name may or may not be the same entity for a given report. |
| Deduplication rule | Whether repeated mentions in one response count once or multiple times; whether the same URL, domain, or citation cluster is deduped | Deduplication separates breadth from repetition. |
| Uncertainty treatment | Run count, confidence interval or range, standard deviation by prompt cluster, and notes on small-sample limits | Without variance, a trend line can mistake normal response volatility for movement. |
Different valid SOV formulas answer different questions. A mention-share formula compares the volume of brand mentions across all brands named. A prompt-incidence formula asks how often your brand appears at least once. A rank-position score asks where the brand appears. A share-of-citation formula asks how often your brand or domain is used as evidence. None of those is the only valid industry formula; the right one depends on the decision the measurement must support.
The canonical AI SOV formula
AI share of voice is the percentage of brand mentions your company receives across AI-generated responses, relative to all brand mentions for your category inside the declared prompt universe. The canonical mention-share formula is:
AI SOV = (your brand mentions / total brand mentions across tracked prompts) x 100
If a prompt set produces 200 counted brand mentions and your brand receives 50 of those mentions, your AI SOV is 25%. That means your brand received one quarter of the counted mention volume in that contract. It does not, by itself, prove recommendation lift, referral traffic, conversion, pipeline, or revenue.
Mention share is the atomic breadth metric. It is complemented by Share of Citation, the atomic depth metric that measures the share of cited evidence assigned to a brand, domain, or source family. Both can roll into a composite AI Visibility Score when the composite weights, inputs, and uncertainty rules are declared. For a broader brand-presence workflow that combines mentions, citations, recommendations, and source absorption, use AI Share of Voice: How to Measure Brand Presence in AI Answers.
Traditional GEO measurement frameworks often focus on attribution and ROI. AI SOV is upstream of that. It tells you whether your brand is present in the answers buyers may read. To connect that presence to business outcomes, measure referral, conversion, pipeline, and revenue separately and test whether changes in AI SOV are associated with changes in those downstream metrics.
A minimal reproducible AI SOV audit
Start with a small contract that is easy to repeat before expanding into a larger tracker.
- Pick one category and one buyer question type. Example: "enterprise PR software" and best-of/comparison prompts.
- Write 20 prompts. Keep the text stable in a versioned prompt library. Label each prompt by category, funnel stage, geography, and buyer role.
- Select provider surfaces. For example: ChatGPT in browsing mode, Perplexity default search, Claude with web access if enabled, and Google AI Overviews as shown in a specified locale. If a product mode cannot cite or browse in the tested state, record that state instead of turning it into a universal platform claim.
- Run each prompt three times. Record the date, time zone, account state, locale, model label visible in the product, response text, mentioned brands, mention rank, linked URLs, and cited URLs.
- Count brands once per response for the core AI SOV score. Use a separate repeated-mention field if you also want intensity. Deduplicate aliases according to the entity rule.
- Calculate totals and ranges. Report mention share by provider, by prompt cluster, and across the full library only if the weighting rule is stated.
- Store screenshots or exports. A later audit should be able to inspect the same evidence, not only the final percentage.
A sample result might read: "In the September 9, 2026 US-English account-neutral audit, 20 best-of and comparison prompts were run three times in each tracked provider mode. AuthorityTech appeared in 18 of 240 prompt-provider runs and received 22 of 310 counted brand mentions, producing a 7.1% mention-share estimate. The provider-level range was 0% to 12.4%, so the next action is provider-specific diagnosis rather than a blended claim."
Do not blend the metrics
The page should not treat every AI visibility number as SOV. Use separate fields:
| Metric | Numerator | Denominator | Best use |
|---|---|---|---|
| AI SOV: mention share | Your counted brand mentions | All counted brand mentions in the same contract | Competitive breadth inside a prompt universe |
| Prompt-level mention incidence | Prompt runs where your brand appears at least once | All prompt runs | Coverage across buyer questions |
| Average rank position | Your brand's ordered position when present | Prompt runs where the brand appears, with a declared missing-rank rule | Prominence within shortlists |
| Share of Citation | Citations or source URLs assigned to your brand/domain/source family | All citations or source URLs in the contract | Evidence depth and source role |
| Link inclusion | Responses that link to your domain or a target source | All responses in the contract | Click opportunity, not mention breadth |
| AI referral traffic | Sessions from identifiable AI referrers | Total site sessions or total AI-referred sessions, depending on the report | Observed traffic, subject to referrer loss and product mode differences |
| Pipeline or revenue association | Qualified opportunities, influenced pipeline, or closed revenue tied to a declared attribution model | The CRM population in the same window | Business impact testing, not SOV itself |
A platform comparison is valid only inside its sampling context. A report can say one provider-mode sample produced more citations or links than another provider-mode sample. It should not say that a provider always behaves that way, or that one product universally has no external links, unless the tested product mode, date, account state, and citation-display conditions are named.
How to build the prompt library
Build the library around the jobs buyers ask AI systems to perform. A practical first version uses four prompt groups:
- Category definition prompts: "What is [category]?" "How does [category] work?"
- Problem prompts: "How do I solve [buyer problem]?" "What should I do when [symptom] happens?"
- Comparison prompts: "[Competitor A] vs [Competitor B] vs [your brand]" or "Compare [category] platforms for [use case]."
- Best-of prompts: "Best [category] tools for [segment]" and "Top [category] platforms for [industry]."
Twenty to fifty prompts can be a useful design range for an initial manual audit, but it is not a universal statistical threshold. The prompt count becomes meaningful only after the universe, repetition policy, variance treatment, and decision tolerance are defined. A 20-prompt diagnostic can find obvious absence; a high-stakes trend report may need more prompts, more repetitions, and provider-level confidence intervals.
Track monthly if the decision is executive reporting. Track weekly or bi-weekly only when you are running an intervention and need faster feedback. Cadence is a design choice, not a validated law.
How to compare providers and product modes
Segment by provider and mode before you aggregate. A useful table has one row per provider-mode pair:
| Provider-mode field | Example value to record |
|---|---|
| Provider/product | ChatGPT, Perplexity, Claude, Gemini, Google AI Overviews, Copilot, Grok |
| Visible model/version | The model label shown to the tester, or "not shown" |
| Mode | Default answer, web browsing, search, deep research, citation-enabled answer, enterprise workspace, or AI Overview SERP |
| Citation surface | Inline citations, source cards, external links, no visible citations, or citations only after expansion |
| Locale/account | US-English signed-out, US-English signed-in, enterprise account, or another declared context |
| Observation window | Exact collection date or rolling window |
| Known limitations | Personalization, inaccessible source cards, missing referrers, inability to rerun a result, or product-mode changes |
Provider segmentation lets you see whether a brand is broadly absent, present only in one product mode, mentioned without links, cited without a positive recommendation, or recommended without a measurable referral path. Those are different problems with different interventions.
How to test interventions without deterministic claims
AI SOV can move after changes to the public evidence environment, but the publication-to-revenue path has many separate steps: publication, crawl or retrieval, entity resolution, mention, citation, recommendation language, link inclusion, referral, conversion, pipeline, and revenue. Do not collapse them into one claim.
Treat each lever as a hypothesis to test:
- Earned media: Hypothesis: independent third-party coverage in sources already visible in your category may improve source availability or citation depth. Test source inclusion and mention outcomes by provider-mode and prompt cluster.
- Entity consistency: Hypothesis: consistent names, product descriptions, authors, and sameAs references may reduce alias fragmentation. Test whether aliases are resolved to the same entity in outputs and citations.
- Owned content: Hypothesis: clear, factual, extractable owned pages may improve answer accuracy when retrieved. Test retrieval, citation, and answer accuracy separately.
- Structured data: Hypothesis: Organization, Article, Person, Product, and FAQ schema can make entity facts easier to parse. Test whether the structured facts appear in retrieved or generated answers; do not assume markup alone changes ranking or recommendation.
- Prompt coverage: Hypothesis: filling missing buyer-question coverage may increase prompt-level mention incidence in the covered cluster. Test the same prompts before and after publication.
- Third-party source network: Hypothesis: consistent independent references across publications, directories, review sites, community discussions, and research pages may create a broader evidence base. Test which source roles are actually cited.
Wikipedia, Wikidata, Google Knowledge Panels, Crunchbase, LinkedIn, and review profiles can be part of the entity-alias record when they exist and are accurate. They should not be described as guaranteed SOV levers. Measure whether the entity is recognized, whether aliases consolidate, and whether the cited evidence supports the answer.
Building an AI Share of Voice measurement stack
A production AI SOV stack needs three layers:
1. Prompt tracking layer
Maintain a versioned prompt library with categories, buyer roles, locale, provider-mode coverage, run count, and collection windows. Keep raw response exports and screenshots where the provider terms and workflow allow it.
2. Mention, citation, and link aggregation
Record brand mentions, prompt-level inclusion, rank position, citations, links, source domains, and sentiment or recommendation language as separate columns. Dedicated AI visibility tools can automate parts of this work, but vendor dashboards still need the same contract fields before their numbers can be compared or reused.
3. Outcome correlation tracking
Track AI-referred traffic where referrers survive, direct-navigation lift where attribution is plausible, and CRM outcomes in a separate model. Match SOV trend lines against pipeline velocity and win rate only as an association until the measurement design supports causal claims. If the business wants a causal read, use holdouts, staggered interventions, or prompt clusters with documented exposure differences.
Machine Relations is the strategic architecture that keeps those layers separate: source authority, entity clarity, citation architecture, distribution, and measurement. The system is useful because each layer can be tested, not because any one layer guarantees the next.
Key takeaways
- AI share of voice is calculated as your brand mentions divided by total brand mentions across a declared prompt universe, multiplied by 100.
- Different valid SOV formulas answer different questions: mention share, prompt incidence, rank position, citation share, and link inclusion should not be blended without a contract.
- The minimum contract names numerator, denominator, sampling unit, prompt universe, provider/model/version/mode, locale/account state, repetition policy, observation window, alias rule, deduplication rule, and uncertainty treatment.
- Platform comparisons are valid only for the tested provider mode, collection window, account state, and locale.
- Prompt-library size and cadence are design decisions. They are not universal thresholds unless the report carries evidence for that population and purpose.
- Earned media, entity consistency, owned content, structured data, third-party sources, and prompt coverage are hypotheses to test, not deterministic levers.
- Publication, retrieval, mention, citation, recommendation, referral, conversion, pipeline, and revenue are separate outcomes.
The Machine Relations system: why AI SOV compounds when it is measured correctly
AI share of voice is useful because it turns machine-mediated discovery into repeatable observation. Every audit shows which prompts, providers, product modes, and source roles are working; every intervention creates a before/after comparison; every comparison improves the next prompt library.
The brands that learn fastest will not be the ones that repeat unsupported benchmark numbers. They will be the ones that keep clean measurement contracts, test interventions against real provider surfaces, and separate visibility from citation, traffic, and revenue until each link is measured.
That operating system is Machine Relations: building relationships with machines the way traditional PR built relationships with journalists, but with the evidence contract made explicit. The formula is not magic. It is consistent third-party validation, structured entity presence, expert-attributed content, prompt coverage, and measurement discipline.
If current AI search visibility is not where it should be, start with a free AI visibility audit to understand exactly where you stand and what is blocking AI share of voice from growing.
How GEO, AEO, and SEO fit within Machine Relations
These disciplines are not competing alternatives — they represent different layers of the same system. Machine Relations is the full architecture that contains each of them:
| Discipline | Optimizes for | Success condition | Scope |
|---|---|---|---|
| SEO | Ranking algorithms | Top 10 position on SERP | Technical + content |
| GEO | Generative AI engines | Cited in AI-generated answers | Content formatting + distribution |
| AEO | Answer boxes / featured snippets | Selected as the direct answer | Structured content |
| Digital PR | Human journalists/editors | Media placement | Outreach + storytelling |
| Machine Relations | AI-mediated discovery systems | Resolved, retrieved, mentioned, cited, or recommended according to the metric being measured | Full system: authority → entity → citation → distribution → measurement |
GEO and AEO are tactics within Layer 4 (Distribution) of the Machine Relations stack. They matter — but they operate on top of source, entity, and measurement layers they cannot replace.
Frequently Asked Questions
What is a good AI share of voice target for a B2B brand?
A good target depends on the category, prompt universe, provider mix, and decision being made. For a first audit, compare your SOV against named competitors inside the same prompt contract and prioritize direction of travel by prompt cluster. Do not treat a vendor benchmark or fixed percentage target as portable unless it uses the same numerator, denominator, provider modes, window, and category.
Which AI platform is most important for share of voice?
The most important platform is the one your buyers use for the job being measured. ChatGPT, Perplexity, Claude, Gemini, Google AI Overviews, Copilot, and Grok can be measured side by side, but each provider-mode row should record model label, browsing or retrieval state, citation surface, locale, account state, and observation date. A blended score is useful only after the weighting rule is declared.
How long does it take to improve AI share of voice?
Use timelines as experiment windows, not promises. A focused intervention can be checked in the next scheduled audit, but a durable trend requires repeated observations across the same prompt library and provider modes. Report the before/after difference with uncertainty and avoid claiming a general 60-90 day or 6-12 month rule unless the evidence matches your category and measurement design.
How is AI share of voice different from traditional share of voice?
Traditional SOV usually compares advertising, media coverage, search visibility, or conversation volume. AI SOV compares brand presence inside AI-generated answers for a declared prompt universe. It can indicate whether a brand is present in machine-mediated research moments, but referral traffic, conversion, pipeline, and revenue require separate measurement.
Can you measure AI share of voice without a dedicated tool?
Yes. A manual audit can use a versioned prompt library, repeated runs, provider-mode labels, entity-alias rules, and a spreadsheet that separates mention share, prompt incidence, rank, citations, links, and notes. Dedicated tools help with scale and consistency, but their outputs still need the same measurement contract before you compare them over time.
Does investing in Wikipedia or knowledge graph presence affect AI share of voice?
Knowledge graph and profile work should be treated as an entity-clarity hypothesis. Accurate Wikipedia, Wikidata, LinkedIn, Crunchbase, Organization schema, and sameAs references may help analysts test whether aliases consolidate and facts are represented consistently. They do not guarantee a mention, citation, recommendation, referral, pipeline, or revenue outcome.