---
title: "AI Brand Rank Is Noise. Inclusion Frequency Is the Signal"
description: "SparkToro and Gumshoe tested 2,961 repeated recommendation prompts across ChatGPT, Claude, and Google AI. List order was unstable; repeated inclusion frequency may be measurable. The study did not identify what causes inclusion."
canonical: https://authoritytech.io/curated/ai-rank-is-noise-consideration-set-is-real-2026
last-updated: 2026-09-09
---

# AI Brand Rank Is Noise. Inclusion Frequency Is the Signal

SparkToro and Gumshoe tested 2,961 repeated recommendation prompts across ChatGPT, Claude, and Google AI. List order was unstable; repeated inclusion frequency may be measurable. The study did not identify what causes inclusion.

Canonical URL: https://authoritytech.io/curated/ai-rank-is-noise-consideration-set-is-real-2026
Published: 2026-03-21
Updated: 2026-09-09
Author: Jaxon Parrott
Tags: Morning Brief, AI Search & Discovery, Newsroom

A brand's position in one AI recommendation list is not a stable rank. In [SparkToro and Gumshoe's January 2026 study](https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/), 600 volunteers ran 12 recommendation prompts through ChatGPT, Claude, and Google AI a combined 2,961 times. The brands returned, their order, and the number of recommendations varied across repeated runs.

For ChatGPT and Google AI, the researchers reported less than a 1-in-100 chance that two of 100 runs would return the same brand list. Getting the same list in the same order was closer to 1 in 1,000. Those figures describe repeated outputs for this prompt set, three systems, and the study's January 2026 collection. They do not establish that every AI product, query, market, or future model behaves identically.

The useful signal was not position. It was **inclusion frequency**: the share of sampled responses in which a brand appeared. That is measurable when the prompt population, system, run count, date, and denominator are disclosed.

## What the Study Measured

The first experiment repeated 12 prompts 60–100 times across ChatGPT, Claude, and Google AI. It measured:

- which brands or products appeared;
- how many recommendations each response returned;
- the order of those recommendations;
- pairwise similarity between responses; and
- the percentage of sampled responses in which each brand appeared.

One ChatGPT example illustrates the difference between frequency and rank. City of Hope appeared in 69 of 71 sampled answers about West Coast cancer hospitals, but it was the first mention in only 25. The study's author explicitly declined to interpret that result as proof that City of Hope was the best hospital or that list position carried a stable meaning.

A second experiment tested prompt variation. Participants wrote 142 prompts expressing the same underlying headphone-shopping intent. The prompts had an average semantic similarity score of 0.081. Across 994 responses generated from those prompts, Bose, Sony, Sennheiser, and Apple appeared in 55–77% of answers.

That result supports a narrower claim: repeated inclusion frequency can remain measurable even when people phrase one intent differently. It does not reveal why those brands appeared, what data each provider used, or which intervention would cause another brand's frequency to rise.

## “Consideration Set” Is an Operational Label

In this context, a consideration set is not a provider-disclosed database or a fixed internal shortlist. It is an operational label for the brands that recur across a defined sample of recommendation responses.

The distinction matters:

- **Rank position** is where a brand appears inside one returned list.
- **Inclusion frequency** is the number of sampled responses that mention the brand divided by all sampled responses in the same panel.
- **Share of Mention** is the brand's mentions divided by all counted brand mentions in the sampled answers.
- **Citation rate** is the number of sampled answers that cite the brand at least once divided by all sampled answers.
- **[Share of Citation](https://machinerelations.ai/glossary/share-of-citation)** is citations attributed to the brand divided by all counted citations in the same sampled answers.

These denominators are not interchangeable. A brand can have high inclusion frequency and low average position, or appear without receiving a citation. None of the five metrics alone establishes recommendation quality, referral traffic, conversion, pipeline, or revenue.

## What Citation Studies Add — and What They Do Not

Citation research can describe the sources attached to generated answers. It does not, by itself, explain which brands enter a recommendation response.

[Ahrefs analyzed ChatGPT's 1,000 most-cited pages in September 2025](https://ahrefs.com/blog/chatgpts-most-cited-pages/). Among the already-cited pages that also ranked in organic search, 65.3% were on domains with Domain Rating 81 or higher; 11.7% were on domains with DR 0–20. That is a distribution within an observed citation inventory. It does not establish Domain Rating as a cause, minimum requirement, provider weighting formula, or mechanism for recommendation inclusion.

[Moz analyzed nearly 40,000 Google AI Mode queries](https://moz.com/blog/ai-mode-citations) across desktop and mobile in the United States and United Kingdom. It found that 88% of AI Mode citation URLs did not exactly match a URL in the organic top 10 for the same query. That result describes citation overlap for Google AI Mode and its sampled queries. It does not establish that organic rank is irrelevant, that off-site coverage controls brand recommendations, or that the same pattern applies to ChatGPT and Claude.

The [GEO benchmark](https://arxiv.org/abs/2311.09735) tested content transformations in a benchmark generative-engine environment and reported visibility gains of up to 40%, with effects varying by domain. It measured benchmark visibility effects, not a universal 30–40% citation-probability lift, current commercial-engine behavior, brand recommendation inclusion, or business outcomes.

Keeping these evidence layers separate prevents a common inference error: joining unstable recommendation order, citation-source composition, organic-search overlap, and benchmark content effects into one claim about what an AI provider “trusts.” The cited studies did not run that joined experiment.

## A Measurement Contract for AI Brand Visibility

A defensible visibility panel should publish enough information for another analyst to reproduce the denominator:

1. **Define the decision being measured.** Recommendation inclusion, brand mention, URL citation, sentiment, and referral are different events.
2. **Define the prompt population.** Record the intents, variants, languages, markets, and whether prompts were human-written or synthetic.
3. **Name the systems and surfaces.** ChatGPT, Claude, Google AI Overview, and Google AI Mode are not one instrument. Model and interface changes matter.
4. **State the collection window.** A panel is a dated sample, not a permanent fact about a brand.
5. **Use enough repeated runs.** SparkToro's study used 60–100 runs per initial prompt; that is evidence for its design, not a universal minimum. Required sample size depends on the variance and decision threshold.
6. **Report the denominator.** “Appeared in 70%” is incomplete unless readers know 70% of which answers, produced from which prompts and systems.
7. **Keep uncertainty visible.** Report counts and intervals where possible. Do not convert one panel's zero into universal exclusion.
8. **Join downstream outcomes separately.** Referral, conversion, pipeline, and revenue require their own observation and identity-matching chain.

The result is a metric you can audit: “Brand A appeared in 62 of 100 responses from this dated prompt-system panel.” It is not a claim that Brand A owns a permanent rank or that the panel has identified the cause.

## What to Do With the Result

Use repeated inclusion frequency to decide where to investigate, not to declare a mechanism.

If a brand's frequency is low, inspect the answer set and source set before prescribing a fix:

- Is the entity named consistently across independent sources?
- Are the sampled prompts asking for a category the brand actually serves?
- Do the systems retrieve or cite pages that mention the brand in the relevant context?
- Are competitors appearing because of current product evidence, reference coverage, reviews, directories, community discussion, or another source class?
- Does the pattern persist across systems, prompt variants, markets, and collection dates?

Each answer is a testable hypothesis. None should be inferred from recommendation frequency alone.

This is where [Machine Relations](https://machinerelations.ai) becomes useful as an operating discipline. It treats brand inclusion, source presence, citation, recommendation, referral, and revenue as connected but separate measurements. The job is to strengthen the public evidence graph, observe how machines use it, and test whether a specific intervention changes a specific outcome.

The strategic correction is simple: stop treating one generated list as a leaderboard. Measure repeated inclusion with an explicit denominator, then investigate the evidence chain that could explain it.

---

See where your brand currently appears across a defined AI prompt panel at [app.authoritytech.io/visibility-audit](https://app.authoritytech.io/visibility-audit).

## Frequently Asked Questions

### Do AI engines rank brands in a fixed order?

SparkToro's January 2026 experiment found highly variable ordering across repeated recommendation prompts in ChatGPT, Claude, and Google AI. That supports treating one response's order as unstable in the sampled systems and prompts. It does not prove that every AI surface lacks every ranking process.

### What is an AI consideration set?

Here, “consideration set” is an operational label for brands that recur across a defined sample of recommendation answers. It is not a provider-disclosed internal list. Measure it as inclusion frequency: sampled answers mentioning the brand divided by all answers in the same panel.

### What controls whether a brand appears in AI recommendations?

The SparkToro study did not test the causes of brand inclusion. Citation inventories, organic-overlap studies, and content benchmarks measure different populations and cannot be combined into a provider mechanism. Diagnose causes with separate source, retrieval, citation, and intervention evidence.

### Can you track AI brand visibility if list order changes?

Yes, if the panel discloses its prompts, systems, run count, date, and denominator. Repeated inclusion frequency can be useful for comparison and trend detection. It should not be reported as a fixed rank, universal market share, or proof of commercial impact.

<!-- AUTO-BACKFILL-LINKS:START -->
## Related Reading
- [PR for AI Search: How Earned Media Drives AI Citation Authority](/industries/pr-for-ai-search)
- [SaaS AI Visibility Strategy: How B2B Brands Get Cited in AI Search](/industries/saas-ai-visibility-strategy)
<!-- AUTO-BACKFILL-LINKS:END -->

---

*Jaxon Parrott is the founder of AuthorityTech and the originator of the [Machine Relations](https://machinerelations.ai/glossary/machine-relations) discipline.*

## Links

- [Curated Index](https://authoritytech.io/curated.md)
- [Home](https://authoritytech.io/index.md)
