How to Get Cited in Gemini: Check Your robots.txt Before Anything Else
Across the 250 most-cited domains in the Machine Relations Index, 59.1% of those Gemini was not observed citing carry Disallow: / on Google-Extended, against 2.9% of those it does. Google documents that token as a grounding control that does not affect Search.
Check your robots.txt first. Across the 250 most-cited domains in the Machine Relations Index, 59.1% of the domains Gemini was not observed citing carry Disallow: / under User-agent: Google-Extended, against 2.9% of the domains it did cite. Google documents that token as a grounding and training control that does not affect Google Search. Most sites that set it kept their Search traffic and gave up Gemini.
That is the single highest-leverage check available, and almost nobody runs it, because Google-Extended looks like an AI crawler and is not one. The rest of this guide covers what the token actually controls, why your server logs cannot confirm whether Gemini ever fetched you, what Gemini in fact cites across 22,179 domains, and the cases where the robots.txt answer is the wrong answer.
Key Takeaways
- Google-Extended is a control token, not a crawler — Google states it "doesn't have a separate HTTP request user agent string" and does not affect Search inclusion or ranking. Blocking it forfeits Gemini grounding while costing nothing in Search, which is why so many publishers set it without noticing the trade.
- The association is strong and it is still only an association — 59.1% against 2.9% across the head of the Index, with counter-cases in both directions that are published below rather than dropped.
- Your access logs cannot answer this question — because the token has no user agent, every request claiming to be Google-Extended is fabricated by definition. In 24 days of edge logs across six sites, 914 requests claimed it.
- Gemini is wide, not narrow — it cited 7,079 distinct domains in the measured window, third of six engines, and reached 452 of the release's 1,222 editorial-media domains, also third.
- The head absences are concentrated in news — 11 of the 23 top-250 domains Gemini was not observed citing are classed as editorial media, and none of the 22 with a readable robots.txt blocks Googlebot outright.
How do I get cited by Gemini?
In order, cheapest first:
- Fetch your own robots.txt and search it for
Google-Extended. If it appears in a group carryingDisallow: /, you have opted out of Gemini grounding. Note that the token is frequently listed inside a block of a dozen AI user agents, so a site can carry it without anyone having decided to block Gemini specifically. - Decide whether that was intentional. It is a legitimate choice. It is only a mistake when it was made to block model training and the cost to answer-time citation was not priced in. The two are governed by the same token.
- Answer the exact question in the first screen of text. Retrievability is necessary, not sufficient — the permission checks below show plenty of domains that permit everything and are still absent.
- Measure at the source level, repeatedly. A citation is not a ranking, a mention, a recommendation, or a referral visit, and the four move independently.
What Google-Extended actually controls
The confusion is structural, and Google's own crawler documentation resolves it in three sentences. From Google's common crawlers reference, read 2026-09-18:
"Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity."
On what it governs, the same page states that Google-Extended manages whether crawled content "may be used for training future generations of Gemini models" and "for grounding (providing content from the Google Search index to the model at prompt time to improve factuality and relevancy)" in Gemini Apps and Grounding with Google Search on Vertex AI.
And on the cost: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."
Read together, those three facts describe a switch that most publishers would want to think about carefully. It bundles training with grounding, so a site that objects to being trained on also stops being quoted at answer time. It is free in Search, so nothing in a traffic dashboard ever registers the decision. And it is enforced without a distinguishable fetch, so no log line marks the moment the site stopped being eligible.
A separate token, Google-CloudVertexBot, covers crawls a site owner requests for their own Vertex AI Agents, and Google notes it "has no effect on Google Search or other products." It is not a substitute and does not restore grounding.
The measurement: head domains, Gemini presence, and robots.txt
The Machine Relations Index observes which sources six answer engines cite for the market's own buyer questions. Release machine_relations_index_public_view_v2.0, generated 2026-09-18, covers 2026-05-10 to 2026-09-18 — 125 days, 15,782 answer runs, 124,397 citation events, 22,179 cited source domains.
Taking the 250 most-cited domains in that release and recording whether Gemini cited each at least once in the window: 227 yes, 23 no. Each of the 250 robots.txt files was then fetched live on 2026-09-18; 227 returned 200. A group counts as a full disallow only when a User-agent: Google-Extended group carries the exact rule Disallow: /.
| Top-250 head domains | Count with readable robots.txt | Name Google-Extended | Carry Disallow: / on it | Block Googlebot entirely |
|---|---|---|---|---|
| Not observed cited by Gemini in the window | 22 | 15 (68.2%) | 13 (59.1%) | 0 |
| Cited by Gemini in the window | 205 | 42 (20.5%) | 6 (2.9%) | 1 |
The thirteen absent domains carrying a full Google-Extended disallow, with their rank in the release: yahoo.com (11), nerdwallet.com (17), techcrunch.com (31), nytimes.com (42), sciencedirect.com (44), cnbc.com (48), instagram.com (56), alibaba.com (59), usnews.com (116), fortune.com (167), investopedia.com (174), consumerreports.org (229), wired.com (250). Any reader can check each one by fetching the file — techcrunch.com/robots.txt, nytimes.com/robots.txt, nerdwallet.com/robots.txt and consumerreports.org/robots.txt each give the token its own group with Disallow: /, while on wired.com/robots.txt and cnbc.com/robots.txt it sits inside a block of a dozen or more AI user agents sharing one Disallow: / — which is how a site ends up blocking Gemini grounding without a decision ever being made about Gemini. Elsevier serves its file only from the apex host: sciencedirect.com/robots.txt returns it, and the www host returns 403.
Not one of the 22 blocks Googlebot outright. These are sites that remain fully available to Google Search and are, across this basket and this window, absent from the engine Google builds on the same index.
The counter-cases, which matter as much
Blocking the token is neither necessary nor sufficient for absence, and the release says so plainly:
- Six head domains carry a full Google-Extended disallow and are cited by Gemini anyway — linkedin.com (3), crunchbase.com (16), amazon.com (20), tracxn.com (64), statista.com (95) and seedtable.com (172) all carry the token under a full disallow and were all cited by Gemini in the window. LinkedIn, third in the whole release, is the sharpest of them. Grounding draws on the Google Search index, and a domain can surface through routes a single token does not close.
- Seven of the absent head domains never name Google-Extended at all — tomsguide.com (71), tech-insider.org (78), impulsec.com (105), projectstartups.com (112), growthlist.co (114), scienceinsights.org (163) and stackmatix.com (168). Their absence has some other cause.
- Two absent domains name the token without closing the door — erpresearch.com (57) gives Google-Extended an explicit
Allow: /and is still absent from Gemini across the window, while being cited by ChatGPT, Claude, Google AI Mode and Perplexity; zoominfo.com (139) lists it in a twelve-agent group that disallows only/p. - One absent domain could not be checked — hrexecutive.com (220) served no readable robots.txt, which is why the denominators above are 22 and 205 rather than 23 and 227.
So this is a correlation, measured on one basket over 125 days, not a mechanism. It is worth acting on because the check costs nothing and the downside of the block is documented by Google itself; it is not worth reporting to anyone as the reason a given site is absent.
How can I tell if I'm being cited in Gemini?
Not from your server logs, and this is the part that sends teams in circles.
Because Google-Extended has no user agent of its own, and grounding serves content from the Google Search index rather than a fresh fetch, there is no request in your logs that corresponds to Gemini citing you. Any log line that appears to show one is fabricated: the string is trivially copied, and Google publishes no agent to match it.
Edge logs across six sites over 2026-08-23 to 2026-09-17 recorded 914 requests presenting Google-Extended as a user agent, across 16 separate days. By Google's own documented contract, none of them could have been what they claimed. That measurement and its method are published in full in the AI-crawler user-agent verification study, which found that no request in the wider set of 59,166 carried a verified identity and that a material share of traffic wearing AI-bot names was credential scanning.
What does work is source-level measurement against a frozen baseline: run the exact prompts you care about on a schedule, record the cited host, the exact placement URL, and the claim the answer leans on, and compare across releases rather than across tools. Per-domain standing in the Index is published at machinerelations.ai/index/domains/ for every one of the 22,179 domains in the release.
Correction, updated September 26, 2026
Three of the six engines the Index monitors had dated collection gaps inside this window, not the two this note reported until today, and the third is Gemini — the engine this page is about. The Index recorded zero citations from Google AI Overviews on any monitored query between August 15 and September 24, 2026 (41 of the last 132 days), zero from Google AI Mode between September 14 and September 24, and zero from Gemini on 22 intermittent days: May 11 to 15, then scattered single days from August 11 to September 16. The two Google lanes share one cause, a defect in the Index's Google collection path that graded a response as a success whether or not it carried content, fixed September 25, 2026. Gemini's intermittency is a separate pattern of single days and its cause is not yet established. ChatGPT, Claude and Perplexity collected throughout; one Claude date was flagged by the same scan and carries exactly one answer run, a scan artifact rather than a gap.
The direction is the whole of this correction, and it runs in this page's favour. A presence count and a distinct-domain count only accumulate, so missing collection days can remove entries from a gapped engine's count but never add them. Every Gemini, Google AI Mode and Google AI Overviews figure on this page is therefore a floor: the true count is at least that and may be higher. A claim that another engine sits below one of the three still holds, because a lower bound above you can only move further above you. A claim that another engine sits above one of them does not, because the floor can rise past it — so every place this page ranked Perplexity or a Google surface above Gemini is withdrawn, while every place it put Claude or ChatGPT below Gemini stands. Gemini against either Google surface is now a gapped arm against a gapped arm and is not stated in either direction. An absence attributed to Gemini is no longer exact either. That matters most for the set this page is built on: the 22 head domains it reports as uncited by Gemini are domains Gemini was not observed citing across the window, and 22 blind days could hide a citation to any of them. The wording throughout has been changed from "never cited" to "not observed citing" for that reason, and the 59.1% figure is a rate computed on an observed-absence set, not a proven-absence set. We are not claiming a direction for that error: a domain wrongly counted as absent would have to be re-read on a clean release to know which way it moves the rate.
What this does not touch is the thesis. This page argues that the common picture of Gemini as a narrow engine is wrong, and a floor can only help that argument: Gemini's breadth is at least what is printed below and may be wider. The robots.txt and crawler-verification findings the page leads on are edge-log measurements taken from our own six sites, not Index citation counts, and are unaffected by any of this.
What Gemini actually cites
The common picture of Gemini as a narrow engine is wrong on this data. Distinct domains cited at least once in the window, with the count cited by that engine alone:
| Engine | Distinct domains cited | Cited by that engine alone | Exclusive share |
|---|---|---|---|
| Perplexity | 9,497 | 4,246 | 44.7% |
| Gemini | 7,079 | 2,837 | 40.1% |
| Claude | 5,138 | 1,431 | 27.9% |
| ChatGPT | 4,877 | 2,708 | 55.5% |
| Google AI Mode † | 7,441 | 3,382 | 45.5% |
| Google AI Overviews † | 2,279 | 365 | 16.0% |
† Gemini, Google AI Mode and Google AI Overviews are floors, not measurements, per the correction above: each cited at least this many domains and may have cited more. Only the Perplexity, Claude and ChatGPT rows are exact. The three floored rows are reported rather than ranked — against each other, or against any row above them.
How to read the exclusive columns. The "cited by that engine alone" and "exclusive share" columns describe tail composition rather than engine breadth. On the Index release of September 21, 2026, 46.2% of the 22,377 cited domains were cited exactly once in the window, and a domain cited once is single-engine by arithmetic, so most of each engine's exclusive pool is the once-cited tail rather than a distinctive source list. Conditioned on evidence the picture inverts: among the 508 domains the Index grades, every one is cited by more than one engine. The distinct-domain counts, the segment-leader agreement counts and the editorial-media counts on this page are frequency-ranked and are the primary measure; the cross-engine agreement study applies the same reading.
Gemini cites at least 45% more distinct domains than ChatGPT, and at least as many as Claude and ChatGPT both — those two arms are exactly measured, Gemini's 7,079 is a floor, so that gap can only widen. Taking the highest-ranked domain in each of the release's 85 published segments — a segment being one subject category paired with one buyer question shape, published only after clearing an evidence floor of 10 observed runs across 7 distinct run dates — and asking which engines cite that leader anywhere in the window: Perplexity 85, Claude 66 and ChatGPT 63 are exact; Gemini is recorded at 78, Google AI Mode at 75 and Google AI Overviews at 73, all three floors. So Gemini agrees with the market's leading source more often than Claude or ChatGPT, and by at least the margin shown. Where it sits relative to Perplexity's 85 this release cannot say: 78 is a lower bound and 85 is inside the range those 22 blind days could account for. The earlier version of this paragraph called Gemini "second widest of the four engines the Index collected continuously" and ranked it behind Perplexity; both claims rested on Gemini being a clean arm, and both are withdrawn.
It also reaches deep into the press: Gemini cited at least 452 of the 1,222 editorial-media domains in the release, well ahead of ChatGPT's exactly measured 328. Perplexity's 516 and the 562 recorded for Google AI Mode are no longer stated as ahead of it — Perplexity is exact but 452 is a floor that could pass it, and AI Mode's own figure is a floor too. What survives is the point the paragraph was making: Gemini cites hundreds of news publishers, which is what makes the head absences above so sharp, and the largest news brands in the head are the ones missing.
At most seven published-segment leaders went uncited by Gemini in the window — an upper bound, because a blind day cannot add a citation but can hide one — and those seven cluster almost entirely in one category: erpresearch.com leads three enterprise-software segments, procuredesk.com and erpfocus.com lead two more, with nerdwallet.com on consumer-finance top lists and paymentsandrisk.com on fintech worth-questions. If you sell enterprise software, Gemini is the engine whose source list least resembles the others in your category. Treat each of those seven as a lead to re-read on the next clean release rather than a settled absence.
How do I increase citations in Gemini?
After the robots.txt check, the work is ordinary and slow:
- Answer the exact query, early. Grounding pulls a passage, not a page. The claim has to survive being lifted out of context, which means a specific sentence with its own numbers and denominator, above the fold.
- Get corroborated somewhere Gemini already reads. In the AI-visibility category, the rank-one source is YouTube on four of the six buyer shapes, Reddit on the worth-questions and HubSpot on head-to-head comparisons — not the trade press. Check what leads in your category before assuming a publication tier.
- Do not treat one engine's source list as the market. Of the 22,179 domains in the release, 14,969 were cited by exactly one engine and only 347 by all six. Optimising for a single engine's pattern optimises for a rounding error.
- Re-measure on each release. Rankings move, and a source that led a segment last month may not lead it now.
Method and limitations
Several third-party studies point the same strategic direction without establishing a Gemini rule, and they are worth reading with their limits attached. Semrush's technical SEO study found many cited URLs have strong technical foundations and frames that as a condition for visibility, not proof that schema wins citations. The Generative Engine Optimization preprint reports broad third-party earned-media patterns and engine differences across ChatGPT, Claude and Gemini, on a different basket to this one. Stacker and Scrunch's wide-republication study reports directional citation lift across an eight-story, five-platform sample, which supports republication as a source-surface intervention and does not prove any bespoke placement causes a named citation.
The measurement on this page has its own boundaries, stated so they can be checked. It covers one basket of the market's buyer questions over 125 days, so a domain absent here may be cited elsewhere for questions the Index does not run. Robots.txt was read on a single day from a single network, and a file that changed last week explains nothing about a window opened in May. Citation and blocking are associated in this head sample; no causal path is demonstrated, and the counter-cases above are the reason to say so. Per-segment standing is read from each domain's published rank in the release rather than re-derived by sorting, because ties make a sort disagree with the publisher. The full release is published as machine-readable data, so every figure here can be recomputed rather than taken on trust.
The check, in one line
curl -s https://yourdomain.com/robots.txt | grep -i -A5 google-extended
If that returns a group containing Disallow: /, you have made a decision about Gemini. Google's documentation says it cost you nothing in Search. This release says the head domains that made it are, with counter-cases, largely the ones Gemini does not cite.
Sources and method. The Google-Extended and Google-CloudVertexBot statements above are quoted from Google's crawler documentation, read on September 18, 2026. The 250-domain head and each domain's Gemini presence were enumerated from the Machine Relations Index release of the same date, and each robots.txt was fetched live from the domain itself; "never cited" means no citation was observed in the monitored Gemini prompts during the release window, not that Gemini cannot cite the domain.
Updated 2026-09-23: figures reflect the current Machine Relations Index methodology; see the MRI methodology and update log.