AI Citation Audit: How to Check Whether AI Engines Actually Cite Your Brand
A systematic framework for auditing whether ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews cite your brand. Covers the three layers of citation measurement, the prompts that matter, and what to do when you discover you are invisible.
An AI citation audit tells you whether ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews mention or cite your brand when buyers ask the questions your company exists to answer. Most brands have never run one. The ones that have usually did it wrong: they searched their own name, got a comforting result, and called it done. That is not an audit. That is confirmation bias with a browser tab.
The real question is not "does AI know my brand?" It is "when a buyer types a category question into an AI engine, does my company appear in the answer, and does the engine link back to my domain as a source?" Those are two different things. The difference between a mention and a citation is the difference between someone saying your name at a dinner party and someone handing a buyer your business card.
I have spent years building measurement infrastructure that tracks citation rates across six AI engines. Here is what I have learned: the gap between brands that think they are visible and brands that actually are is enormous. This piece is the audit framework I wish someone had published when we started.
Why Most AI Citation Checks Are Useless
The standard advice is simple: open ChatGPT, type your brand name, and see what comes back. Every guide starts there. And every guide that stops there is giving you noise instead of signal.
Here is the problem. AI answers are non-deterministic. The same question produces different sources an hour later. A single query in a single engine on a single afternoon tells you almost nothing about your actual citation rate. It tells you what happened once. You need what happens consistently.
The second problem is worse. Searching your brand name tests whether the engine knows you exist. That is the easiest test to pass and the least useful to pass it. The queries that matter are the ones where a buyer does not type your name at all. "Best AI PR agencies." "How to get cited by AI search engines." "What is the difference between earned media and paid media for AI visibility." If you are absent from those answers, you have a pipeline-affecting citation problem, regardless of how well the engine describes you when asked directly.
The Three Layers of Citation Measurement
Not every check is equal. There are three distinct layers, and each one answers a different question.
Layer 1: Spot-check. You open an AI engine, type a query, and look at the result. This is a thermometer reading. It tells you the temperature right now. It does not tell you whether you have a fever pattern. Useful for a first look. Dangerous as a strategy input.
Layer 2: Structured audit. You build a library of 20 to 50 non-branded queries, run them across multiple engines in clean sessions, and log presence, citation depth, and sentiment. SuperData SEO outlines a version of this that covers eight engines. Everything PR publishes a 35-prompt framework organized across six query types: brand identity, category, comparative, buyer-intent, founder, and geographic. This is the minimum viable measurement.
Layer 3: Citation rate. You measure how often AI engines cite your domain across enough observations and enough days that the number is stable. A citation rate is not "ChatGPT mentioned me 3 out of 5 times today." It is a rate calculated from at least 10 observations across at least 7 distinct dates, with a confidence grade reflecting the volume of evidence. This is the layer where you stop guessing and start managing. FreeCodeCamp's analysis of 7 sites found visibility-to-citation gaps ranging from 25 to 95 points, proving that being known and being cited are structurally different problems.
Most brands never reach Layer 3. The ones that do stop asking "are we visible?" and start asking "what is our citation rate in this segment, and how does it compare to the competitor who is winning?"
How to Run a Structured AI Citation Audit
Here is the framework. It takes 60 to 90 minutes for the first pass and 30 minutes for each monthly refresh.
Step 1: Build your query library. Write 25 to 35 queries a buyer would actually type. Not branded queries. Category queries. Problem queries. Comparison queries. Decision queries. Everything PR's framework organizes these into six types, which is a clean structure. The questions your sales team hears on calls are the best starting material.
Step 2: Choose your engines. Test at minimum: ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews. If you want completeness, add Copilot and Google AI Mode. Each engine selects sources differently. Topic Intelligence's platform-specific analysis documents the sourcing divergence: ChatGPT relies heavily on Bing indexing and Wikipedia, Perplexity weights Reddit and fresh content published within 30 days, and Gemini favors schema-rich owned content. A brand optimized for one engine's sourcing logic can be structurally absent from another.
Step 3: Run clean sessions. Use fresh, non-personalized sessions. Logged-in accounts carry history that skews results. Private/incognito mode. No prior conversation context. This matters because personalization changes what the engine surfaces.
Step 4: Score every result. For each query and engine combination, record:
| Signal | What to Track |
|---|---|
| Presence | Is your brand named in the answer text? |
| Citation | Is your domain listed as a source with a link? |
| Position | Where in the answer does your brand appear? |
| Sentiment | Is the description accurate and positive? |
| Competitor | Which competitors appear instead of you? |
SuperData SEO's scoring weights these: a linked citation earns 2 points, a mention earns 1, absence earns 0. Apply a sentiment modifier: +1 for accurate positive, 0 for neutral, -1 for inaccurate.
Step 5: Calculate your citation share. Divide the number of queries where you were cited by the total number of queries tested, per engine. This is your raw citation share. For a more granular view, Topic Intelligence recommends weighted scoring: 3 points for a definitive mention (primary recommendation), 1 point for a supporting mention, and -2 for a negative citation. Brands scoring below 40 out of 175 possible points across five engines face what practitioners call a pipeline-affecting gap.
What the Audit Results Actually Mean
The numbers tell you four things.
Engine divergence. Your citation rate will vary dramatically across engines. Research shows that Google AI Mode and AI Overviews select sources by different logic: market databases and research firms earn 1.6 to 1.9 times more citations in AI Mode, while wire services and industry publishers index higher in AI Overviews. Your competitive set changes depending on which engine you measure.
The mention-citation gap. Being mentioned is not the same as being cited. A mention means the model names your brand in the answer text; a citation means your domain appears among the sources. Many brands are mentioned without citation, which means the engine learned about you from someone else's content. You get the name recognition without the traffic.
Category vs. brand visibility. If you score well on brand queries but poorly on category queries, you have an awareness problem that AI is exposing, not creating. The engine knows who you are. It does not consider you a definitive source on the problems you solve. Trysight's tracking framework calls these "indirect citations": cases where the AI describes your offering without naming your brand, signaling missed attribution opportunities that only a structured audit reveals.
Competitive displacement. The "cited instead of you" list from your audit is your actual AI search competitor set. These are often different from your traditional search competitors. A competitor you have never tracked in SEO may be dominating your AI citation landscape because they have source architecture you do not.
The Mention vs. Citation Distinction and Why It Changes Everything
This deserves its own section because most brands miss it entirely.
When an AI engine mentions your brand, it is repeating something it learned during training or retrieved from someone else's page. When it cites you, it is pointing the user to your domain as a source. The first is a fact about the model's knowledge. The second is a fact about your source architecture.
HubSpot's tracking framework makes this distinction central. GeoReady classifies results into four tiers: Strong (cited consistently), Cited (present but inconsistent), Mentioned only (known from third-party pages), and Invisible (neither mentioned nor cited).
If your audit shows "Mentioned only" across most queries, the fix is not more brand awareness. You are already known. The fix is source architecture: making your content directly retrievable, clearly structured, and independently citable. Peer-reviewed research from Aggarwal et al. found that writing for justification rather than keyword density can boost AI visibility by up to 40%. That means your domain needs to be the place where the answer lives, with clearly attributed sources and quotable statements, not a brand that other sources write about.
How Often to Run the Audit
A one-time audit is a snapshot. AI engines update their indexes and models continuously. GeoReady recommends weekly measurement as "a sensible floor for a category you care about." SuperData SEO recommends monthly for high-stakes brands, quarterly for others.
Here is what I recommend based on what we measure:
- Weekly: Track your top 10 category queries across your 3 most important engines. This takes 15 minutes and catches drops before they compound.
- Monthly: Run the full 35-query audit across all engines. This is your citation rate baseline.
- After every major publish: Spot-check the queries your new content should influence. Trysight's tracking data shows a 2 to 4 week lag between content publication and citation impact changes. If you published a definitive guide on X and the engines are not citing it within that window, you have a source architecture problem, not a patience problem.
Presence AI notes that citation rates can shift 12 points in a single week when model updates roll out. If you are not measuring regularly, you will not know whether a drop is temporary volatility or a structural change.
Tools for Scaling Beyond Manual Audits
Manual auditing works at 25 to 50 queries. Beyond that, you need tooling.
Indexly outlines a setup that takes 2 to 4 hours initially and requires 30 to 60 minutes weekly after that. AnswerManiac recommends tool-based tracking when your prompt list exceeds 50 queries. The tooling landscape is maturing fast: HubSpot, Semrush, and several AI-native platforms now offer some version of citation tracking.
The critical requirement for any tool is that it separates mentions from citations, tracks results over time (not just point-in-time), and measures across multiple engines. A tool that only tests ChatGPT gives you one-sixth of the picture.
At the systematic level, what matters is not any single tool but the discipline of measuring citation rates with enough observations, across enough days, to produce a stable number. A rate from 3 spot-checks is a guess. A rate from 10 observations across 7 days is a signal you can act on.
What to Fix When the Audit Shows You Are Invisible
An audit that ends with "we are invisible" and no action plan is a waste of the 90 minutes you spent. Here is the decision tree.
If you are invisible across all engines and all query types: Your domain is not being retrieved by any AI engine as a source. The problem is foundational: weak crawlability, no structured content, or no independent third-party sources that reference your domain. FreeCodeCamp's diagnostic thresholds classify visibility below 20% as a distribution issue requiring brand mentions in communities, forums, and independent publications before content optimization will have any effect. Fix retrieval before fixing content.
If you are mentioned but not cited: AI engines know about you from other people's content, but they are not pulling from your domain directly. Build source architecture: clear answer-first content, structured data, and pages that directly address the queries where you want to be cited.
If you are cited on one engine but invisible on others: Each engine has different source-selection logic. Research on AI Mode vs. AI Overviews shows that source-type preference varies significantly. Diversify your source architecture: earn citations from independent publications, build comparison-ready structured content, and ensure your domain is independently crawlable by all major AI bots.
If you are cited on brand queries but not category queries: You have recognition without authority in your category. Peer-reviewed research shows that AI systems demonstrate significant authority bias, preferring third-party editorial coverage over brand-owned content. The fix is content that owns category-level questions, supported by entity chains: independent sources across multiple domains that all reference your brand in the context of the category problem.
The Audit Is the Starting Line, Not the Finish
Running an AI citation audit tells you where you stand. It does not fix where you stand. The brands that treat the audit as a recurring measurement discipline, not a one-time curiosity, are the ones that end up in the answers.
The shift is already here. AI Overviews can reduce organic CTR by 15 to 34% for queries that trigger a summary. LoudPixel's tracking data shows AI citations convert at 14.2% versus 2.8% from Google organic, roughly 5 times the value per click. The buyers who used to click ten blue links now read one synthesized answer and follow the sources it names. If your brand is not one of those sources, you are not losing a ranking. You are losing the conversation entirely.
This is Machine Relations. Not SEO with a new acronym. A fundamentally different measurement surface that requires a fundamentally different audit discipline. The question is not whether you should run this audit. The question is what you are going to do when the results come back and your biggest competitor is in every answer where you are absent.
Go run the audit. Build the query library. Test five engines. Score the results. Do it again next month. The brands that measure this systematically are the ones the machines learn to cite. The ones that do not are the ones wondering why their pipeline dried up while their SEO dashboard still looked green.
FAQ
How long does an AI citation audit take?
The first structured audit takes 60 to 90 minutes: 20 minutes to build the query library, 30 to 40 minutes to run queries across five engines, and 10 to 20 minutes to score results. Monthly refreshes take about 30 minutes. Weekly spot-checks of your top 10 queries take 15 minutes.
Which AI engines should I test?
At minimum: ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews. These cover the primary surfaces where B2B buyers ask questions. SuperData SEO recommends testing eight engines including Copilot, You.com, and Google AI Mode for completeness. Start with five and expand based on where your buyers actually spend time.
What is the difference between a mention and a citation in AI search?
A mention means the AI engine names your brand in its answer text. A citation means your domain appears among the linked sources the engine used to generate that answer. Mentions indicate the engine knows about you. Citations indicate it trusts your domain as a primary source. Only citations drive traffic back to your site.
How many queries should I include in my audit?
Start with 25 to 35 non-branded queries organized by type: category, comparison, problem-solution, buyer-intent, and geographic. Everything PR's 35-prompt framework generates 175 data points when run across five engines. Scale to tool-based tracking when your list exceeds 50 queries.
What score means my brand has a citation problem?
Brands scoring below 40 out of 175 possible points across five engines face pipeline-affecting gaps. If your brand is absent from more than half of category-level queries, you are effectively invisible to buyers using AI search for purchase research. The threshold depends on your industry, but below 25% citation share on category queries is a clear signal to act.