Afternoon BriefAI Search & Discovery

5WPR Says 15 Websites Own 68% of AI Citations. In Your Category, the Top 10 Hold 7-24%

5WPR's 680-million-citation study puts 68% with 15 websites, pooled across every topic. Cut by category, the Machine Relations Index puts the top ten at 7.0-24.2% of a category's citations. Here's the founder playbook for cross-engine citation architecture.

Jaxon Parrott
Jaxon ParrottMay 16, 2026

Most founders I talk to are still optimizing for Google rankings. Meanwhile, a small set of trusted domains controls an outsized share of every answer that ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews produce.

The best-known version of that finding, 15 websites holding 68% of citations, comes from the 5WPR AI Platform Citation Source Index 2026 — a consolidated analysis of 680 million AI citations across five major AI engines, pooled across every topic at once. The Machine Relations Index measures the number a founder can plan against: cut by category across 23,280 domains, the ten most-cited domains in a category hold 7.0 to 24.2 percent of that category's citations, and reaching half a category's citation volume takes 49 to 304 distinct domains depending on the category — six times apart at the extremes, and never as small as 15 (category breakdown and method). The concentration is real; size a strategy against the category number, not the pooled 15-site figure.

If your brand isn't present on the surfaces those engines trust, you don't exist at the moment your buyer asks the question.

The Citation Map Is Not What You Think

Here's what the data actually shows about where AI engines source their answers:

EnginePrimary Citation SourcesBehavior
ChatGPTWikipedia, Reddit, Forbes, Business InsiderSelective — 7-8 citations per response from high-authority sources
PerplexityPrimary research, NIH/PubMed, named B2B authorityBroad — 20-22 inline citations, rewards original data
GeminiFirst-party documentation, Knowledge Graph entities36-40 citations, leans on official sources
ClaudeNYT, The Atlantic, The New Yorker, The EconomistOnly 36% of journalism citations from the past 12 months
Google AI OverviewsYouTube (200x advantage), Reddit, community contentInline links next to claims since May 6 update

Reddit alone accounts for roughly 40% of all citations across LLMs. Wikipedia captures 26-48% of ChatGPT's top-10 share. YouTube holds a 200x citation advantage over every other video source in Google AI Overviews.

This isn't a level playing field. It's a concentrated oligarchy of trusted sources — and each AI engine has a different set of preferences.

The Volatility Problem Nobody Plans For

Here's what should scare you: ChatGPT's Reddit citation share fell from 60% to 10% in six weeks in late 2025. One parameter change. PR Newswire, Forbes, and Medium absorbed the displaced share.

Citation visibility is now measured in weeks, not years. Your entire AI visibility can collapse overnight because of a single upstream change you don't control and won't be warned about.

Semrush's AI Visibility Study confirmed that AI citations change 40-60% month over month. If you're treating AI visibility as a set-it-and-forget-it SEO exercise, you're building on sand.

Why Traditional PR Fails This Test

A traditional PR agency lands you a Forbes article. That works for ChatGPT — Forbes is in its top citation sources. But it barely registers on Perplexity, which rewards primary research and named B2B authority. And it does almost nothing for Claude, which preferentially cites legacy editorial outlets like The New York Times and The Atlantic.

One placement on one surface covers one engine. You need presence across the citation architectures that all five engines trust.

This is what I built Machine Relations to solve. Not press releases. Not media lists. Source-architecture strategy that maps your brand's citation presence to the specific surfaces each AI engine reads, verifies, and recommends.

What the Google May 6 Update Confirms

Google just shipped five structural changes to AI Mode and AI Overviews — the biggest citation-surface update since AI Overviews launched in 2024. The key changes:

  • Inline links next to claims — not grouped at the bottom anymore. The unit of optimization is now the passage, not the page.
  • Community Perspectives — Reddit threads, forums, and expert blogs now get quoted with attribution inside AI answers.
  • Branded web mentions correlate 0.664 with AI appearances — nearly 3x higher than backlinks (0.218), per DemandSignals.

Meanwhile, Google's March 2026 core update elevated first-party brand sites while punishing aggregators. YouTube lost 567 visibility points. Reddit lost 64. Instagram lost 48. X lost 46. First-party brand sites gained.

The message is clear: Google wants original sources cited, not platforms that aggregate and write about them. If you're the actual authority — and you've built the citation architecture to prove it — you win on both traditional search and AI surfaces.

The Founder Playbook: Cross-Engine Citation Architecture

Here's what I tell every founder who asks me how to become visible in AI answers:

1. Audit your presence across the top 15 citation sources first. Not your website's SEO. Your brand's presence on the surfaces AI actually cites: Reddit, Wikipedia, YouTube, LinkedIn, G2, industry publications, primary research databases. If you're invisible on these, your website ranking is irrelevant.

2. Build per-engine citation strategies. ChatGPT needs Wikipedia entity clarity and Forbes/Business Insider mentions. Perplexity needs primary data and named authority. Claude needs quality editorial coverage. Google AI needs structured content with passage-level extractability. A single "content strategy" won't cover this.

3. Treat Wikipedia as infrastructure. Not as a one-time PR win. Your entity page is the foundation that AI systems use to resolve your brand's identity. If you don't have one, or it's thin — you're starting at a structural disadvantage.

4. Plan for citation volatility. Build presence across multiple surfaces in each engine's preference set so that when one surface shifts (and it will), your visibility doesn't collapse. Diversification isn't optional anymore.

5. Measure share of citation, not just rankings. Search your top 10 buyer-intent queries across ChatGPT, Perplexity, Google AI Mode, Claude, and Gemini. Document which brands get cited. If you're not among them, that's your competitive reality — regardless of what Google Search Console says about your organic position.

The Bottom Line

The AI citation economy is winner-take-most, but the winners' circle is category-specific, not a fixed list of 15 sites — our own measurement puts the top ten's share of a category's citations at 7.0 to 24.2 percent, not a flat 68 (category breakdown and method). Each engine has different trust signals. The landscape shifts in weeks. And AI-referred traffic converts 23x higher than traditional organic visitors.

This is not a content optimization problem. It's a source-architecture problem. You either engineer your brand's presence across the citation surfaces that matter for your category — or you watch from outside while the domains that already dominate your category's answers capture the attention that used to come through search.

I've been saying this since I coined Machine Relations: the game isn't ranking pages. The game is becoming a source AI engines trust enough to cite. The 5WPR study put 680 million data points behind the general shape of that claim; our own category-level index is the more useful number for sizing a specific plan.

Updated 2026-09-23: figures reflect the current Machine Relations Index methodology; see the MRI methodology and update log.


FAQ

How many domains control most AI citations in 2026? The 5WPR AI Platform Citation Source Index 2026, analyzing 680 million citations across ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews, found that 15 domains capture 68% of all AI citation share pooled across every topic at once. That is not the number to plan against: cut by category, our own measurement finds the ten most-cited domains hold 7.0 to 24.2 percent of a category's citations, and reaching half a category's citations takes 49 to 304 distinct domains depending on the category.

What is the most-cited source across AI engines? Reddit is the number-one cited source across every major AI engine, accounting for approximately 40% of all AI citations. Wikipedia is the second most-cited, with 26-48% of ChatGPT's top-10 citation share. The measured unit is source composition in the cited sample. Boundary to preserve: this does not establish a universal source-selection mechanism, citation causation, a guaranteed future citation, recommendation lift, pipeline, revenue, or any other business outcome.

Do different AI engines cite different sources? Yes. ChatGPT favors Wikipedia, Reddit, Forbes, and Business Insider. Perplexity rewards primary research and named B2B authority. Claude leans toward legacy editorial outlets like The New York Times and The Atlantic. Google AI Overviews favors YouTube and community content.

What is Machine Relations? Machine Relations is the discipline of earning AI citations and recommendations for a brand by making that brand legible, retrievable, and credible inside AI-driven discovery. It replaces traditional PR for the AI era by treating citation architecture — not media impressions — as the primary visibility metric.

Additional source context