Stop Building Your AI Citation Pitch List Around 15 Sites
The '15 websites own 68% of AI citations' number is still circulating, including on one of our own pages. Our own 23,280-domain measurement shows the real number a pitch list needs, and it is category-specific, not 15.
A number has been circulating in citation-strategy content since May: 15 websites own 68 percent of all AI citations. It runs across at least two dozen marketing and PR sites this quarter, and — until this week — it ran on this site too, in the lead sentence of a briefing about citation strategy for founders.
Our own measurement says something different. It says the number was never one number to begin with.
Where the 68 percent came from
The figure traces to the 5WPR AI Platform Citation Source Index, a synthesis of six earlier studies covering roughly 680 million citations across ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews, announced via PR Newswire, reported by Everything PR and repeated by outlets including GPT Melo and Perea. Its top 15 names are dominated by platforms, not publishers: Reddit at roughly 40 percent on its own, Wikipedia, YouTube, LinkedIn and Amazon alongside Forbes, Business Insider, TechRadar, Reuters and the New York Times.
None of that is wrong as a description of where citations concentrate on the open web across every topic mixed together. The problem is what a founder or comms team does with it next: reads "15 sites, 68 percent" and concludes a pitch list of 15 outlets covers most of what an AI engine will cite for their category. Reddit, Wikipedia and YouTube are not outlets a PR team pitches, and a single cross-topic index tells a B2B security buyer's category nothing about which 15 sources actually matter for that category's questions.
Two other trackers this quarter measured the same phenomenon at smaller scale and got smaller, category-specific numbers: Attrifast's vertical study found six domains capture roughly 71 percent of citation slots in healthcare versus 28 percent for the top six in SaaS — the concentration ratio itself moves by category, not just the winners. Gadex's B2B benchmark put ChatGPT's ten most-cited domains at 34.4 percent of its citations and Perplexity's ten at 21.7 percent, in the same sample, on the same engines. Three different studies, three different concentration figures, because concentration is a property of the category and the engine, not a portfolio-wide constant.
What our own data shows
Machine Relations' public index measures 23,280 cited domains across six engines over a 130-day window (release mri_score_v2.0, generated 2026-09-23, 16,475 answer runs, machinerelations.ai/index). Cut by category rather than pooled across the whole index:
| Measure | Low end | High end |
|---|---|---|
| Top-10 domains' share of a category's citations | 7.0% (Industrial) | 24.2% (Family Software) |
| Distinct domains needed to reach half a category's citations | 49 (Family Software) | 304 (Legacy News Topics) |
The ten most-cited domains in a category hold between 7.0 and 24.2 percent of that category's citations, and reaching half a category's citation volume takes between 49 and 304 distinct domains depending on the category (confidence grade B–C on category aggregates, evidence floor 10 observed runs across 7 run dates). That is the opposite of "15 sites, 68 percent" holding for every category at once — it is a range that moves six-fold by category, and even at its most concentrated measured category (Family Software, 49 domains to reach half) the list is more than triple the 15-site figure being repeated.
There is a second reason the 68 percent figure overstates what a pitch list controls. Of the citations Machine Relations' index can classify into a source type, editorial publications — the kind of outlet a PR pitch actually targets — hold 1,222 of 22,179 cited domains and 13,309 of 110,877 domain-run citations, led by Medium at 6.1 percent of that class. 60.5 percent of all citation volume in the index sits in an unclassified layer the nine defined source-role classes don't cover. A pitch list built around a known top-15 platform list targets a small, already-saturated slice of a much larger and more fragmented citation surface.
The pitch-list-sizing gap this exposes
Media-list guidance outside AI search already sizes lists well past 15 for real reach. PR Lab's build guide puts a working campaign list at 25 to 100 relevant contacts and a master list at 50 to 500-plus, segmented by beat and tier. Contentgrip's media-relations framework and Cision's pitch guidance both start from defining the story and audience before naming a single outlet, precisely because the right list is a function of category and story, not a fixed universe.
That is the same conclusion the citation data reaches from a different direction. A list sized to the 68-percent figure is sized for the wrong population — a cross-topic mix of platforms most brands can't pitch — when the number that should drive list size is the category's own domain-count-to-half figure, which runs from roughly 49 to over 300 depending on category.
What to do with this
- Pull your category's own concentration figure before setting a pitch-list target, not the portfolio-wide "15 sites" number. A Family Software brand and a Legacy News-adjacent brand need list sizes six times apart.
- Split the list by source-role class, not just outlet size. Editorial publications are one class among nine measured; a list built only from recognizable media names misses vendor-owned, market-database and community sources that carry real citation share in most categories.
- Treat platform-level figures (Reddit, Wikipedia, YouTube) as a different problem from pitch-list sizing. Platform presence and earned-media placement are different mechanisms; a study measuring both together makes the earned-media slice look smaller than the list it should actually inform.
Correction note
An earlier AuthorityTech briefing, 15 Websites Own 68% of All AI Citations — What Founders Should Do About It, opened with this figure sourced to the same 5WPR index. That figure describes platform-level concentration across the whole open web, not the category-specific, publisher-focused number a pitch list should be sized against — the measured range above is the correction, and this page supersedes that claim as of 2026-09-23.
Sources and method
Category concentration figures are read from the Machine Relations public index, release mri_score_v2.0 generated 2026-09-23 (window 2026-05-10 to 2026-09-23, 130 observed days, 23,280 cited domains, 129,264 source events, 16,475 answer runs, six engines: ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, Perplexity; evidence floor 10 observed runs across 7 run dates), machine-readable at machinerelations.ai/data/machine-relations-index.json. The 68-percent figure and its component numbers are read from the 5WPR AI Platform Citation Source Index as reported by Everything PR, cross-checked against independent trackers from Otterly, Attrifast and Gadex, which each report different concentration ratios for different samples — itself evidence that concentration is not a single portfolio-wide constant.
FAQ
Is the "15 websites own 68 percent of AI citations" number fake? No — it is a real reading of a specific cross-topic synthesis of citations to platforms like Reddit, Wikipedia and YouTube alongside a handful of major publishers. It is not, however, a usable target for sizing a category-specific earned-media pitch list, which is how it has been repeated.
How many publishers should actually be on an AI citation pitch list? Size it to the category, not a fixed number. Machine Relations' index puts the distinct-domain count needed to reach half a category's citations at 49 to 304 depending on category — closer to standard PR media-list sizing guidance (25 to 100-plus contacts) than to a flat 15.
Does this mean citation concentration isn't real? No. The top 10 domains in every measured category still hold a disproportionate 7.0 to 24.2 percent of that category's citations. What changes is the list size needed to move beyond that top tier into the half that decides most category-level visibility.