AI Citation Volatility: Why Brand Mentions Fluctuate
The 25 most-cited domains in AI answers were absent on 42.4% of observed days across a 125-day, six-engine window. Citation volatility measured on our own instrument, plus the engine-breadth hypothesis it falsifies.
Across the 125-day window ending 2026-09-18, the 25 most-cited domains in the Machine Relations Index were absent from the cited source set on 1,325 of 3,125 observed domain-days — 42.4%. The median domain in that group was cited on 69 of 125 days. Being one of the largest sources in AI answers does not mean being present in them on any given day.
AI citation volatility is the rate at which an answer engine changes the URLs or domains it cites for the same measurement setup: frozen prompt, engine, locale, account state, and scan interval. It is not the same thing as brand mention stability, recommendation inclusion, cited-host share, or exact-URL citation. Those are adjacent outcomes, and each needs its own denominator.
What 125 Days of Our Own Measurement Shows
The numbers below come from the Machine Relations Index, release mri_score_v2.0+2026-09-18+8fa38e54dd0a: 15,782 answer runs across ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, and Perplexity, producing 124,397 source events and 110,877 domain-run citations over 22,179 cited source domains, observed daily from 2026-05-10 to 2026-09-18.
One field in that release answers the durability question directly. For each domain it records days_cited over days_observed: the share of observed days on which the domain was cited at least once anywhere in the panel. Read it as presence, not as prompt-level churn. It does not say a specific prompt kept a specific URL from one scan to the next; it says whether the domain was in the answer set that day at all. That is the weaker, more robust claim, and it is the one the data supports.
Here is the whole-Index top 25 by citation volume, with presence beside it.
| # | Domain | Source role | Citation rate | Runs cited | Days cited (of 125) | Presence | Engines | Avg. position |
|---|---|---|---|---|---|---|---|---|
| 1 | reddit.com | Community and social | 12.75% | 2,012 | 120 | 0.960 | 4 | 9.9 |
| 2 | youtube.com | Search or media platform | 9.07% | 1,431 | 81 | 0.648 | 6 | 8.2 |
| 3 | linkedin.com | Community and social | 6.81% | 1,075 | 116 | 0.928 | 6 | 10.0 |
| 4 | medium.com | Editorial publication | 5.11% | 807 | 108 | 0.864 | 6 | 9.2 |
| 5 | forbes.com | Editorial publication | 4.11% | 649 | 104 | 0.832 | 6 | 7.0 |
| 6 | nih.gov | Academic and government | 3.07% | 485 | 88 | 0.704 | 6 | 8.5 |
| 7 | gartner.com | Analyst and consulting | 2.88% | 454 | 91 | 0.728 | 6 | 8.4 |
| 8 | techradar.com | Editorial publication | 2.84% | 448 | 99 | 0.792 | 4 | 5.4 |
| 9 | g2.com | Market and company database | 2.80% | 442 | 91 | 0.728 | 6 | 7.9 |
| 10 | arxiv.org | Academic and government | 2.60% | 410 | 69 | 0.552 | 6 | 6.0 |
| 11 | yahoo.com | Editorial publication | 1.94% | 306 | 67 | 0.536 | 5 | 8.7 |
| 12 | microsoft.com | Vendor-owned | 1.93% | 305 | 78 | 0.624 | 6 | 8.2 |
| 13 | ibm.com | Vendor-owned | 1.93% | 304 | 66 | 0.528 | 6 | 7.1 |
| 14 | landbase.com | Vendor-owned | 1.91% | 301 | 49 | 0.392 | 6 | 6.7 |
| 15 | healthline.com | Other observed source | 1.88% | 297 | 15 | 0.120 | 6 | 6.4 |
| 16 | crunchbase.com | Market and company database | 1.72% | 271 | 43 | 0.344 | 6 | 5.5 |
| 17 | nerdwallet.com | Editorial publication | 1.60% | 253 | 31 | 0.248 | 5 | 6.0 |
| 18 | substack.com | Community and social | 1.58% | 249 | 76 | 0.608 | 5 | 8.9 |
| 19 | techtarget.com | Editorial publication | 1.50% | 237 | 79 | 0.632 | 6 | 6.1 |
| 20 | amazon.com | Editorial publication | 1.50% | 236 | 66 | 0.528 | 6 | 9.3 |
| 21 | paloaltonetworks.com | Vendor-owned | 1.37% | 216 | 48 | 0.384 | 6 | 7.7 |
| 22 | dev.to | Community and social | 1.32% | 209 | 56 | 0.448 | 6 | 7.7 |
| 23 | facebook.com | Community and social | 1.29% | 203 | 52 | 0.416 | 4 | 11.0 |
| 24 | sentinelone.com | Vendor-owned | 1.27% | 200 | 52 | 0.416 | 6 | 7.6 |
| 25 | deloitte.com | Analyst and consulting | 1.24% | 196 | 55 | 0.440 | 6 | 8.2 |
Three things fall out of that table.
Volume rank and presence rank are different rankings. They correlate (Spearman 0.78 across the 25), then come apart in exactly the places that matter to a brand. healthline.com is 15th by citation volume, with 297 cited runs, and last of the 25 by presence: 15 days of 125. That is a burst, not a position. crunchbase.com and nerdwallet.com each fall seven places between the two rankings; youtube.com, second by volume, is ninth by presence. Nine of the 25 were cited on fewer than half the observed days.
Reaching rank 1 separates nothing. All 25 of these domains have held the top cited position at least once. Best-ever position is therefore a useless discriminator in the head of the index. What varies is average position, which runs from 5.4 for techradar.com to 11.0 for facebook.com, and presence, which runs from 0.120 to 0.960. A tool that reports your best rank is reporting the field's most crowded statistic.
The one source class a brand fully controls is the least durable one in the head. Grouping the 25 by source role and averaging presence: community and social 0.672 (n=5), search or media platform 0.648 (n=1), editorial publication 0.633 (n=7), academic and government 0.628 (n=2), analyst and consulting 0.584 (n=2), market and company database 0.536 (n=2), vendor-owned 0.469 (n=5), other observed source 0.120 (n=1). Vendor-owned domains average 0.469 presence against 0.603 for everything else in the group, and four of the five sit below 0.53; microsoft.com is the only one above. Small cells, so treat the ordering as a signal rather than a ranking — but the direction is the uncomfortable one for owned-media budgets.
Does Engine Breadth Predict Persistence? The Index Tests It
A common hypothesis holds that a brand present across multiple independent source roles may become less dependent on any one page, one publisher, or one engine. The release above can test the engine leg of it directly, because engine_breadth is recorded per domain.
It does not hold in the head of the index.
Across the 25 domains, the Spearman correlation between engine count and presence is −0.07 — no relationship in either direction. The 19 domains cited by all six engines average 0.571 presence. The six domains cited by fewer than six engines average 0.593, marginally higher. The most persistent domain in the entire Index, reddit.com at 120 of 125 days, is cited by four engines, not six. Seven of the nineteen six-engine domains were absent on more than half the observed days, healthline.com among them at all six engines and 15 days.
Two limits on that result, both real. Breadth is near-universal in this group — 19 of 25 sit at all six engines — so the head of the index is a weak place to separate breadth's effect from anything else, and a test across the full 22,179 domains could land differently. And this tests only the engine leg. Whether presence across multiple independent publishers or source roles predicts persistence is a different measurement on a different axis, and it remains open.
What the result does close is the easy version of the argument. "Get cited by every engine" is not, on this evidence, a durability strategy. Breadth of engine coverage and persistence of citation are separate outcomes, and buying the first does not deliver the second.
What AI Citation Volatility Actually Measures
Similarweb defines citation volatility as change in the specific URLs and domains cited by AI engines for a given prompt between consecutive measurement scans. A high-volatility prompt has many cited sources replaced from one scan to the next. A low-volatility prompt has a more persistent citation set.
Use that definition at the prompt level. A brand can have high AI visibility and unstable cited sources, or modest visibility with a more stable citation set. Neither state proves durable authority by itself. The measurement question is narrower: when you repeat the same prompt under the same conditions, how much of the cited source set survives into the next scan?
This also keeps causality honest. A drop in cited URLs can follow a model update, a retrieval-policy change, competitor publishing, third-party source edits, account personalization, locale changes, or ordinary run-to-run variation. The observation alone does not prove which driver caused the movement. Our presence figures above are subject to the same caution: they describe what the panel observed, not why a domain left the answer set on a given day.
The Outside Evidence, With Its Actual Unit
Public research on citation churn is real but heterogeneous. It measures different populations, engines, intervals, and units, and the numbers do not collapse into one constant.
| Source | Engines / sample | Period or interval | Unit and finding | How to use it |
|---|---|---|---|---|
| Knecht Strategies | Internal tracking plus unspecified public-tool benchmarks for mid-sized B2B brands | 2025–2026 practitioner monitoring; public method not fully disclosed | 40–60% of citations typically changing over 30 days in its operating context | Practitioner risk signal. Do not treat as a cross-industry constant without a disclosed sample and formula. |
| Trakkr Study 009 | 10K brands; page displays 7 AI models while the narrative says 8 | 10-month daily brand-query study; last updated March 30, 2026 | 70.5% of URLs cited only once; 29 days from peak to half | Citation-lifespan warning with a disclosed public inconsistency in model count and withheld source-domain examples. |
| Similarweb | Prompt-level AI Brand Visibility examples | Consecutive scans | Citation volatility for URLs/domains; Chanel perfume visibility example moves 86% to 14% | Defines the metric and shows movement. The example does not establish the cause of the movement. |
| Sielinski / IQRush arXiv:2607.10341 | 30 platform-topic tests across Gemini, SearchGPT, and Perplexity Search | One collection period; up to 125 citation-bearing responses per platform-topic test | 33–94 citation-bearing answers for the two stopping conditions among 27 converged tests; 3 of 30 unresolved at 125 | Justifies repeated sampling, confidence intervals, and platform-topic segmentation. It is not a fixed query budget for every market. |
| SparkToro / Gumshoe | 2,961 responses from 600 volunteers, 12 prompts, ChatGPT, Claude, and Google AI surfaces | Volunteer runs under usual settings | Recommendation lists rarely repeated exactly, and ordered lists almost never repeated exactly | Evidence about recommendation-list variability, not URL-citation half-life or citation ranking stability. |
| SISTRIX | 82,619 qualified prompts and 1,548,213 snapshots across Germany, USA, UK, Italy, Spain, and France | 17 weekly observations from December 17, 2025 to April 8, 2026 | Domain-level churn differs by platform: AI Mode around 56% weekly, ChatGPT Search up to 74%, AI Overviews much steadier for many prompts | Platform-specific churn evidence. Do not merge it into one generic durability cause. |
| Stacker / Scrunch | 3M+ citation events, 120K+ non-network domains, 8 industries, 6 AI platforms | 26-week window from September 2025 to March 2026 | Non-network domains had roughly 4.5-week citation half-life overall; OpenAI non-network domains had a 3.4-week half-life; Stacker network domains were longer in that cohort | Cohort-level persistence benchmark. It does not prove that every earned-media placement creates durability. |
| Semrush | 230K prompts across ChatGPT Search, Google AI Mode, and Perplexity; top 25 cited domains weekly | July 14 to October 12, 2025 | Reddit appeared in close to 60% of ChatGPT responses in early August and around 10% by mid-September | Shows a sharp platform-specific citation mix change. Semrush investigated possible causes; it did not establish a retrieval-change cause. |
The common takeaway across all of them is methodological rather than numerical: freeze the measurement setup, repeat the prompt panel, and report uncertainty before deciding whether a change is meaningful. Our own figures sit alongside these as one more population with one more unit — a daily six-engine panel measured at domain-day presence — not as a replacement constant.
Why One Snapshot Is Not Enough
A single scan can be useful as an inventory: it tells you which answer, citations, and source roles appeared under one set of conditions. It should not be used as a trend, durability, or causal claim.
The presence column above is the concrete reason. A single scan of the panel on a day when healthline.com happened to be cited would place it among the largest sources in AI answers. A scan on any of the other 110 days would not find it at all. Both scans would be accurate. Neither would be a durability measurement.
Presence is also the unit that survives when clicks do not. Pew Research Center analysed the browsing activity of 900 U.S. adults who agreed to share it, and found that 58% ran at least one Google search in March 2025 that produced an AI summary, that users were less likely to click a result link on pages carrying an AI summary than on pages without one, and that they very rarely clicked the sources cited in the summary. On that behaviour, being in the cited set on the day a buyer asks is the outcome to measure, and a referral-traffic series is a poor proxy for it.
The IQRush preprint is the cleanest caution because it separates two requirements: rank stability and structural sufficiency. In 27 of 30 platform-topic combinations, both conditions were met somewhere between 33 and 94 citation-bearing answers. Three SearchGPT combinations did not satisfy the sufficiency condition even after 125 citation-bearing responses. That does not mean every topic needs exactly 94 submitted queries. It means sample adequacy depends on the platform, topic, citation yield, dispersion, and the comparison you intend to make.
A practical measurement contract should therefore report:
- the exact prompt panel and any prompt variants;
- engine, model surface, locale, account state, and personalization state;
- scan timestamp and comparison interval;
- cited URL, cited host, brand mention, and recommendation outcomes separately;
- presence across observed days, not only rate within cited runs;
- sample size, citation-bearing response count, and missing or zero-citation behavior;
- confidence intervals or a clear "not enough data yet" state.
Candidate Drivers to Investigate
When citation volatility appears, investigate drivers instead of assigning cause from the chart alone.
1. Model and retrieval changes
Semrush's 2025 study shows how quickly a top source mix can move inside one platform: Reddit was present in close to 60% of ChatGPT responses in early August and around 10% by mid-September. Semrush discusses the timing around Google's removal of the num=100 parameter and other possible explanations, then states that the data does not prove why the shifts happened. Treat platform events as hypotheses to test against your own frozen panel.
Google's own documentation supplies a mechanism that predicts day-to-day movement without any platform event at all. Google Search Central states that AI Overviews and AI Mode "may use a 'query fan-out' technique — issuing multiple related searches across subtopics and data sources — to develop a response", and that while responses are generated the models identify more supporting pages, producing a wider and more diverse set of links than a classic web search. The same page directs publishers to Google's fundamental SEO best practices as the guidance that applies to appearing in AI Overviews and AI Mode. If the subquery set is generated per response, a domain's presence on a given day is partly a property of that fan-out rather than of the domain, which is what a presence figure like 0.120 for a large source describes.
2. Platform-specific source selection
SISTRIX shows that Google AI Overviews, Google AI Mode, and ChatGPT Search had different domain-level churn patterns over the same 17-week study. Similarweb also frames volatility at the prompt level, not as one all-engine average. Measuring one engine can tell you something about that engine under that setup; it does not automatically transfer to every other engine. The engine-breadth result above is the counterpart at the domain level: presence at all six engines did not make a domain more persistent in our window.
The platforms themselves document retrieval as a per-engine concern. OpenAI publishes its crawler and user-agent documentation for the bots behind ChatGPT search and training, and Perplexity documents its own crawlers with separate WAF and IP-range guidance for site operators. Two engines, two retrieval stacks, two sets of operator controls. Averaging your visibility across them produces a number that belongs to no engine.
3. Source-ecosystem and competitor movement
Similarweb's Chanel example shows a visibility movement from 86% to 14% for a perfume prompt context. The page discusses competitive activity and source instability as possible drivers, but the example itself does not prove that a competitor improvement caused the drop or that a brand's own pages were irrelevant. Use it as a prompt to inspect the replacement sources, not as a counterfactual.
Measurement Workflow for Citation Volatility
- Freeze the panel. Define prompts, engine surfaces, geography, language, account state, and scan cadence before collecting data.
- Separate outcomes. Track exact-URL citations, cited-host share, brand mentions, recommendation inclusion, and answer sentiment as different fields.
- Record presence separately from rate. Log the days on which the domain was cited at all, not only its citation rate within cited runs. Those two numbers rank the same set differently, as the table above shows.
- Repeat runs. Use multiple citation-bearing answers per prompt-engine segment when the decision depends on small differences.
- Report uncertainty. Show ranges, confidence intervals, or a "needs more observations" flag instead of clean point estimates when the sample is thin.
- Compare like with like. Analyze ChatGPT Search, Perplexity, Google AI Mode, and Google AI Overviews separately before presenting any portfolio-level summary.
- Set alert thresholds from your own baseline. A high-risk buyer prompt may warrant checks every few days; a low-risk informational prompt may not. Cadence is a decision-risk choice, not a universal research requirement.
- Investigate before attributing. When a citation drops, inspect replacement sources, query variants, platform timing, source freshness, and whether the movement exceeds normal run variation.
This workflow complements broader AI visibility measurement. If you need the full metric taxonomy, use the AI visibility measurement guide and the AI Visibility Score definition as adjacent references, but keep volatility as its own stability measure. For placing your own citation rate against the measured distribution, use what a good AI citation rate actually is. To check any specific domain's own presence and segment ranks against this release, look it up in the Index.
What This Means for Machine Relations
Citation volatility is one reason Machine Relations treats the public evidence graph around a brand as operational infrastructure rather than a campaign. The measurement above sharpens what that infrastructure can and cannot be expected to do.
It can reasonably be expected to create more places to observe the brand, which is what makes a change detectable at all. On this release it cannot be expected to produce persistence through engine coverage alone: breadth across all six engines showed no relationship to presence across the 25 most-cited domains, and the least persistent source role in that group was the vendor-owned one.
The remaining legs of the hypothesis — independence across publishers, and across source roles — are still worth proving or falsifying, and the way to do it is unchanged: a frozen baseline, dated interventions, repeated post-change windows, and relevant comparison groups. Breadth alone is not durability, and we now have our own number saying so rather than someone else's.
FAQ
How often do AI citations change?
It depends on the engine, prompt, source type, and measurement interval. On our own daily six-engine panel over 125 days, the 25 most-cited domains in the Machine Relations Index were absent on 42.4% of observed domain-days, with a median of 69 days present out of 125 and a range from 15 days to 120. Externally, Knecht reports 40–60% monthly turnover from practitioner tracking for mid-sized B2B contexts, while SISTRIX reports weekly domain-level churn that differs across Google AI Mode, Google AI Overviews, and ChatGPT Search. Report the exact interval and denominator before comparing any of these numbers to each other.
Why did my AI visibility drop suddenly?
A sudden drop can come from normal run variation, model or retrieval changes, competitor publishing, third-party source updates, geography, account state, or source-category shifts. The drop itself is an observation, not a diagnosis. Given that nine of the 25 largest sources in AI answers are absent on most observed days, a single-day disappearance is closer to the base rate than to an event. Compare repeated runs and inspect the replacement sources before assigning cause.
Can I make AI citations more stable?
Treat every answer here as a hypothesis to measure rather than a guarantee. One common answer now has evidence against it: getting cited by more engines did not predict more consistent presence in our window, with a Spearman correlation of −0.07 across the 25 most-cited domains. The most persistent domain in the Index was cited by four engines, not six. Improving source quality, making content easier to parse, and earning independent corroboration remain reasonable practices, but no combination of them guarantees that a specific URL or publisher will remain cited on a given day.
Is a high AI citation rate the same as being consistently cited?
No, and the two rank the same domains differently. Citation rate is the share of observed runs in which a domain was cited; presence is the share of observed days on which it was cited at all. Across the whole-Index top 25 the two rank orders correlate at Spearman 0.78 and then diverge sharply: healthline.com ranks 15th by rate and 25th by presence, and youtube.com ranks 2nd by rate and 9th by presence. If your dashboard reports one number, find out which one it is.
How many times should I query an AI engine to get reliable visibility data?
Use the number of citation-bearing responses required for the decision, not a universal query count. In arXiv:2607.10341, 27 of 30 platform-topic tests met both stopping conditions between 33 and 94 citation-bearing answers, while 3 did not converge by 125. Smaller changes, lower citation yield, or tightly clustered competitors require more evidence.