Start your AI visibility journey with as little as $49/mo
Every AI Engine Has Its Own Citation DNA: What 120,000 Citations Reveal About ChatGPT, Gemini, Perplexity & Claude (Reddit Deep-Dive)

Every AI Engine Has Its Own Citation DNA: What 120,000 Citations Reveal About ChatGPT, Gemini, Perplexity & Claude (Reddit Deep-Dive)

Aug 11, 2026
|
Akshat

Every AI engine cites a different internet.

Ask ChatGPT, Gemini, Perplexity, and Claude the same question in the same category and you will get four answers pulled from four almost non-overlapping source pools. In one recent 120,000-citation analysis we ran for a single B2B category, the top-cited domain on Perplexity accounted for 26% of that engine’s total citations — and the same domain accounted for 7.9% on ChatGPT. Some sources that dominate ChatGPT’s citation list (thousands of citations) are cited effectively zero times by Claude on the same query set. YouTube and Reddit are Perplexity plays. Patents and analyst reports are Claude plays. Manufacturer-official domains are ChatGPT plays.

Different engines. Different citation DNA. Same category, same buyers, same prompts.

Here is the Reddit-style breakdown of engine citation trends, what the 120K-citation dataset actually shows, and what to do about it if you are running an AEO or GEO program in 2026.

AI engine citation DNA hero

The Study: 119,939 Citations, 4 Engines, 30 Days

The dataset behind this post is an anonymized 30-day citation trend report from a B2B category with a single dominant vendor and 15–20 credible category competitors. It covers 119,939 total citations analyzed across four AI engines — ChatGPT, Gemini, Perplexity, and Claude — for the same tracked prompt set.

We stripped out any category identifiers and vendor names before publishing. What is generalizable is not the specific domains — it is the pattern of engine-specific source affinity, which we have now seen repeat across every B2B category we have measured.

Here is what the split looks like at the engine level:

EngineCitations analyzed (30d)Share of total
Gemini49,83641.5%
Perplexity39,66433.1%
Claude15,71813.1%
ChatGPT14,72112.3%

The first surprise for most teams: Perplexity produces roughly 2.7× the visible citations of ChatGPT for the same prompt set. Perplexity’s answer format is citation-heavy by design — 5–15 sources per answer is typical. ChatGPT with search enabled surfaces 2–6 citations per answer, and often produces answers with no visible citations at all. Citation share (percentages), not raw counts, is the fair comparison across engines.

The Core Finding: Every Engine Has Its Own Citation DNA

The single most important pattern in the data is that the same category is being answered from radically different source pools depending on which engine you ask.

Here is a simplified view of the top-cited domain type per engine, from the anonymized dataset:

EngineWhat it cites mostExample category type of top sources
ChatGPTCategory-canonical OEMOfficial manufacturer & technical reference sites
GeminiVendor + vertical pubsOwned domain + specialist industry publications
PerplexityUGC + video + ownedOwned domain + Reddit, LinkedIn, YouTube, aggregators
ClaudeOwned + patents + analystOwned domain + USPTO, market-research reports

If you flatten this into an “invest here to influence engine X” heuristic:

  • To move ChatGPT: publish on category-canonical technical/manufacturer domains and long-form authoritative reference pages.
  • To move Gemini: publish on vertical industry publications and get indexed thoroughly through Google’s existing crawl.
  • To move Perplexity: invest in Reddit, LinkedIn thought leadership, YouTube, and aggregator/marketplace listings.
  • To move Claude: file patents, get cited in analyst reports (Markets and Markets, Mordor Intelligence, Yole), and get listed in specialized directories.

Same content strategy will not move all four engines. This is the operating truth that most brand teams still budget against as if AEO were a single channel.

Engine citation DNA comparison

Skew: The Metric That Actually Directs Investment

Raw citation share is useful, but the metric that directs where to spend is skew — how disproportionately one engine cites a domain compared to the others.

A domain with 1.5× skew on Perplexity is cited (share-wise) 1.5 times more by Perplexity than by the other three engines on average. Skews of 5×–10× are common. In the analyzed dataset, several domains showed skews above 10×, meaning one engine treats that source as authoritative for the category while the others largely ignore it.

Some of the most instructive skew patterns from the anonymized dataset:

PatternSkewed towardPractical read
Category-canonical manufacturer domain #1ChatGPTDeep technical reference content wins on ChatGPT
Category-canonical manufacturer domain #2ChatGPTReinforces the same pattern
Category-canonical manufacturer domain #3ChatGPTThree in a row — ChatGPT strongly prefers OEM domains
Video platformPerplexityYouTube investment moves Perplexity more than others
Vertical industry publicationGeminiGoogle index still routes Gemini toward specialist pubs
Component-distributor / aggregatorPerplexityAggregator listing quality is a Perplexity lever
Patent database (USPTO)ClaudeClaude reaches for primary technical/legal sources
Analyst / market-research reportsClaudePaid analyst placements translate into Claude visibility
RedditPerplexityReddit sentiment is a direct Perplexity input
LinkedInPerplexityPerplexity treats LinkedIn as a professional-context source

Skew is the single most actionable citation metric because it maps directly to which content investments move which engines. A brand that runs a source-affinity matrix once a quarter and re-allocates budget against it will out-invest a brand that treats AEO as a single undifferentiated channel.

Engine-By-Engine Breakdown

Here is the practical read on each engine’s citation behavior from the study — with the specific tactics that follow from it.

ChatGPT: Category-Canonical Wins

Total citations analyzed: 14,721.

ChatGPT’s top three cited domains in the dataset were all category-canonical manufacturer / OEM websites, each holding a 9–12% citation share. Together the top three accounted for ~33% of all ChatGPT citations for the query set — a heavy concentration in a small number of authoritative sources.

The brand’s own domain came in fourth at 7.9%. Aggregators, marketplaces, video platforms, and social sources were effectively absent from ChatGPT’s citation list in the analyzed category.

Tactical read for AEO on ChatGPT:

  • Owned technical documentation and deep reference pages are the primary lever.
  • Category-canonical trade sites and OEM domains are the earned-media targets (where you can influence via joint content, whitepapers, or authoritative placements).
  • Social, aggregator, and video investment is not the ChatGPT play. It moves other engines instead.

Gemini: Owned + Vertical Publications

Total citations analyzed: 49,836 — the largest citation volume of any engine in the study.

The brand’s own domain led Gemini at 11.7% citation share. Below the brand, Gemini spread its citations across a wider pool than ChatGPT — vertical industry publications, category-canonical manufacturers, video, and aggregator/distributor sources all featured meaningfully (each in the 1.5–3.2% range).

The pattern reflects Gemini’s tight integration with Google’s index: the same signals that make a site rank well in Google Search increase its likelihood of being cited by Gemini. Vertical trade publications with strong Google authority disproportionately show up.

Tactical read for AEO on Gemini:

  • Owned domain SEO fundamentals still matter — the brand’s own content is Gemini’s #1 citation source.
  • Vertical industry publications are the highest-leverage earned-media targets.
  • YouTube is a meaningful (though not dominant) Gemini lever — a video content investment on Google’s own platform gets rewarded.
  • Aggregator and distributor listings are a soft lever for categories with those middle-layer players.

Perplexity: The UGC and Video Engine

Total citations analyzed: 39,664.

Perplexity’s citation pattern is the most distinctive of any engine. The brand’s own domain dominated at 26.0% share — a single-domain concentration that reflects Perplexity’s tendency to lean hard into a canonical source when it can identify one. Below the brand, the top cited sources included:

  • Category-canonical manufacturer (5.5%)
  • YouTube (4.9%)
  • Component distributor (3.1%)
  • Another category-canonical manufacturer (3.1%)
  • Vertical pub / industry news
  • Reddit (2.0%)
  • LinkedIn (1.9%)

The presence of YouTube, Reddit, and LinkedIn in the top 10 makes Perplexity the UGC + video engine for this category. None of those three sources appeared with any weight in ChatGPT’s or Claude’s citation lists.

Tactical read for AEO on Perplexity:

  • Reddit sentiment and Reddit thread presence are direct inputs into Perplexity answers. Ignoring Reddit as an AEO channel systematically underinvests in Perplexity visibility.
  • LinkedIn thought leadership and executive presence is a Perplexity signal in a way it is not for other engines.
  • YouTube is a strong Perplexity lever — video tutorials, product demos, and category explainers on YouTube get cited.
  • Aggregator and distributor listings matter — Perplexity treats marketplace/aggregator pages as authoritative comparison sources.

Claude: The Primary-Source Engine

Total citations analyzed: 15,718.

Claude’s citation profile was the most conservative and most “primary-source” of the four. The brand’s own domain led at 14.7%, with category-canonical manufacturers in the top three. Then something interesting happened — the USPTO patent database (image-ppubs.uspto.gov) appeared as Claude’s 4th most-cited domain at 3.2% share, followed by:

  • Analyst / market-research site (2.2%)
  • Specialized industry directory (1.9%)
  • Category-canonical manufacturer (1.4%)
  • Component distributor (1.3%)
  • Analyst research firm — Yole Group (1.2%)
  • Analyst research firm — Mordor Intelligence (1.1%)

Two analyst firms and a patent database in the top 10. This is Claude’s signature citation pattern: when in doubt, reach for a primary or authoritative analytical source.

Tactical read for AEO on Claude:

  • Patent portfolios translate more directly into Claude visibility than into any other engine’s visibility.
  • Paid analyst-report placements (Markets and Markets, Mordor Intelligence, Gartner, Yole, IDC) are a Claude lever.
  • Specialized industry directories with rich per-entity data are worth being listed in.
  • Social and video investment does not move Claude at all — assume it is not part of Claude’s citation universe.

Source affinity matrix by engine

What Never Gets Cited (And What That Tells You)

Zeroes matter as much as high shares. In the analyzed dataset:

  • ChatGPT never cited YouTube, Reddit, LinkedIn, or Mouser for the tracked category.
  • Claude never cited YouTube or Reddit meaningfully (below 0.1% share).
  • Some vertical trade publications and industry directories were cited only by Gemini and never by ChatGPT.

Zeroes are the strongest signal in a source-affinity matrix. If your content strategy for ChatGPT relies on Reddit or YouTube distribution, the data says your investment cannot reach ChatGPT through that channel at all — you have to route through owned domain and category-canonical earned media instead.

The step-by-step process to turn per-engine citation data into a distribution priority queue:

1. Run a per-engine citation audit. Take 200–500 representative buyer prompts for your category. Run them against ChatGPT (with search enabled), Gemini, Perplexity, and Claude on a scheduled cadence — weekly is a reasonable starting point. Parse every visible citation URL and aggregate by root domain, per engine.

2. Build a source-affinity matrix. Rows = top 20–30 cited domains. Columns = engines. Cells = citation share, with skew highlighted (favored = green ≥1.5×, avoided = red ≤0.5×, dash = never cites).

3. Classify each cited domain. Bucket every domain into:

  • Owned — you already control it.
  • Earnable — a publication, forum, directory, aggregator, or platform where you can realistically place content or earn a mention in 3–6 months.
  • Unreachable — patents, government, competitor domains, opaque sources.

4. Weight by engine priority. Multiply each earnable domain’s citation share by the strategic weight of the engines it appears on. Perplexity-strong domains matter more if your buyer research skews Perplexity. Claude-strong domains matter more for enterprise/technical categories.

5. Rank and allocate. The ranked list of weighted earnable domains is your distribution priority queue for the quarter. Publish content, earn placements, submit to directories, or invest in analyst relations in that order.

6. Re-measure monthly. Track how the citation-share distribution shifts as your investments land. This closes the loop between AEO investment and measured citation impact — see our GEO playbook for the full framework and the prompt-monitoring guide for how to size the tracked prompt set.

Why This Changes How Brands Budget for AEO

Most 2026 AEO programs still budget as if AI search were a single channel. The 120K-citation dataset says otherwise.

  • A ChatGPT-first budget buys deep owned technical content and OEM/reference-site placements. It does not buy Reddit, YouTube, or LinkedIn distribution.
  • A Perplexity-first budget buys Reddit sentiment monitoring, LinkedIn thought leadership, YouTube content, and aggregator listing quality. It does not require heavy analyst-firm spend.
  • A Claude-first budget buys patent filings, analyst-report placements, and specialized directory listings. It does not buy social distribution.
  • A Gemini-first budget buys traditional Google SEO fundamentals on the owned domain and vertical-industry-publication placement.

Very few brands have a single-engine budget. Most need a weighted mix across all four engines based on their buyer base — and the source-affinity matrix is the tool that turns that mix into a concrete quarterly distribution plan.

That is what citation tracking is for. Not measurement for measurement’s sake — directing capital toward the channels that actually move the answer on the engines your buyers use.

Sanbi.ai’s citation analytics runs each brand’s tracked prompt set against ChatGPT, Gemini, Perplexity, and Claude on a scheduled cadence, parses every visible citation, and aggregates into:

  • Top cited domains per engine (share + raw count, exportable)
  • Domain × engine matrix with green/red skew highlighting
  • Affinity list — domains disproportionately cited by a single engine (≥1.5× skew), which drive engine-specific investment decisions
  • Time-series — per-domain, per-engine citation share over time, with alerts on statistically significant shifts (a new integration ships, an index refreshes, a competitor gains ground)
  • Competitive citation share — how your citation share compares to named competitors on each engine

The visualizations in this article — top-cited-domains table per engine, domain × engine matrix — are drawn from Sanbi’s PDF export format. Brand teams use them for quarterly board reporting and AEO budget defense.

Full mechanics in our complete AEO guide, the multi-engine tracking guide, and the prompt-monitoring guide.

The One-Line Takeaway

Every AI engine cites a different internet — and until you can see which one each engine is reading from, your AEO budget is guessing.

Citation tracking turns that guess into a directed investment. Start with a source-affinity matrix. Weight the earnable domains by engine priority. Publish and place in that order. Re-measure. Repeat.

That is the whole playbook.


Related reading:

Frequently Asked Questions

What is LLM citation tracking and why does it matter in 2026?

LLM citation tracking is the discipline of measuring which specific source domains an AI answer engine (ChatGPT, Gemini, Perplexity, Claude, Copilot, DeepSeek) links to, quotes, or paraphrases when it generates an answer. It matters because each engine draws from a different underlying source pool — the same query returns citations to completely different sites across engines. Without citation tracking, a brand cannot see where its authority signals are landing, which competitors are being cited instead, or which publications and platforms it needs to influence to change the answer. In 2026, citation tracking sits alongside prompt monitoring as the two core measurements of AI visibility.

Do ChatGPT, Gemini, Perplexity, and Claude cite the same sources?

No — and it is not close. In a 120,000-citation study across a single B2B category, the top-cited domain on Perplexity received 26% share of that engine's citations, while on ChatGPT the same domain received 7.9%. Some domains that dominate ChatGPT (with 5,000× total citations) barely register on Claude. Other domains — like YouTube and Reddit — are heavily cited by Perplexity but effectively never cited by ChatGPT or Claude in the same category. Each engine has its own citation DNA driven by different training corpora, different retrieval architectures, and different content-format preferences.

Which AI engine cites Reddit the most?

In the citation dataset analyzed, Perplexity was by far the largest citer of Reddit (2.0% of Perplexity's total citations vs. 0.3% on Gemini and effectively 0% on ChatGPT and Claude for the same query set). This tracks with the wider pattern: Perplexity's live-retrieval architecture treats Reddit threads as high-authority sources for opinion, recommendation, and 'is X worth it' style prompts. For brands, that means Reddit sentiment is a direct input into Perplexity answers in a way it is not for ChatGPT or Claude — and Reddit visibility strategy must be scoped by engine, not treated as a universal AEO tactic.

Which AI engine cites YouTube the most?

Perplexity again — YouTube was the third-most-cited domain on Perplexity (4.9% share, 1,947× total) versus 1.7% on Gemini and zero citations from ChatGPT and Claude in the analyzed dataset. Perplexity's answer format frequently embeds YouTube results directly, especially for demo, tutorial, and product-review queries. Gemini also cites YouTube meaningfully (owner advantage — YouTube is Google's), while ChatGPT and Claude in most default configurations do not cite video content at all. This has direct implications for video content strategy: a YouTube-heavy investment influences Perplexity and Gemini disproportionately.

Which AI engine is most likely to cite patents and academic sources?

Claude — in the analyzed dataset, image-ppubs.uspto.gov (the USPTO patent database) was the 4th most-cited domain on Claude at 3.2% share, and market-research sites like marketsandmarkets.com and mordorintelligence.com featured heavily. This reflects Claude's tendency toward more conservative, source-grounded answers on technical and analytical queries. ChatGPT and Perplexity cited these sources far less. For B2B brands in technical categories, this means patent portfolios and paid market-research reports translate into Claude visibility more directly than into ChatGPT or Perplexity visibility.

What is citation share and how is it measured?

Citation share is the percentage of an engine's total outgoing citations (in a given tracked query set and time window) that go to a specific domain. If Perplexity generated 39,664 citations across a query set and 10,332 of those pointed to domain X, then domain X has 26.0% citation share on Perplexity for that query set. Citation share is measured by running a defined prompt set against each engine on a scheduled cadence, parsing the citations, aggregating them by root domain, and dividing by total citations per engine. Sanbi.ai's citation analytics reports citation share per engine, per domain, per time window.

What is 'skew' in AI citation analysis?

Skew (or 'affinity') measures how disproportionately one engine cites a domain compared to the others. A domain with 1.5× skew on ChatGPT is cited 1.5 times more (share-wise) by ChatGPT than by the other engines on average. Skews of 5×–10× are common. In the analyzed dataset, ti.com had a 7,179× total citation count that was dominated by ChatGPT — meaning ChatGPT treated ti.com as an authoritative source for the category while other engines diluted their citations across a broader source pool. Skew is the single most actionable citation metric because it tells you which content investments influence which engine.

How do I find out which domains each AI engine cites for my brand?

The manual approach: run 20–50 representative buyer prompts against ChatGPT (with browsing enabled), Gemini, Perplexity, and Claude; copy every citation URL from each answer; aggregate by root domain; count and compute share per engine. This works for a one-off audit but is not sustainable. Automated citation tracking platforms — Sanbi.ai being one — run the same prompt set against multiple engines on a schedule, parse citations at scale, deduplicate, aggregate by domain, and report affinity and skew across engines in a single dashboard. For any brand tracking more than 25 prompts, automation is a hard requirement.

What is brand citation tracking in ChatGPT?

Brand citation tracking in ChatGPT measures how frequently ChatGPT cites your brand's own domain (owned media), how frequently it cites third-party sources that mention your brand (earned media), and how the two combine into total share of voice for your brand vs. competitors. The measurement runs a tracked prompt set against ChatGPT (with browsing/search enabled) on a scheduled cadence and parses the citation footer for domain hits. In our 120K-citation study, the target brand's own domain accounted for 7.9% of ChatGPT citations for its category — below the top three third-party sources, illustrating a common pattern where owned media is not the dominant source of AI-answer visibility.

How does citation behavior differ between ChatGPT with browsing and ChatGPT without browsing?

ChatGPT with browsing (search enabled) behaves closer to Perplexity — it fetches live web results, cites URLs, and reflects recent content. ChatGPT without browsing draws from training data and typically does not surface visible URL citations at all — brand mentions still occur, but they are unattributed. Citation tracking of ChatGPT should be scoped to the browsing / search-enabled surface where visible citations exist. For unattributed brand mentions inside pure-training-data responses, sentiment and mention tracking (not citation tracking) is the correct instrument.

Do LinkedIn posts influence AI search citations?

Yes, on Perplexity specifically. In the analyzed dataset LinkedIn accounted for 1.9% of Perplexity's citations (756× total) for the tracked query set, while receiving effectively zero citations from ChatGPT and Claude and only 2 total from Gemini. This pattern is consistent across other B2B categories we have seen: Perplexity treats LinkedIn as a legitimate professional-context source, while other engines do not. For a B2B brand, that means an aggressive LinkedIn thought-leadership investment moves Perplexity visibility more than it moves visibility on other engines.

How do industry-specific publications compare to social platforms as AI citation sources?

It depends on the engine. Industry-specific publications (trade press, analyst reports, category-specific news sites) dominate on Gemini and Claude, which lean toward domain-authoritative sources. Social platforms (Reddit, LinkedIn, YouTube) dominate on Perplexity, which weights user-generated and multimedia sources heavily. ChatGPT sits between the two but skews toward official manufacturer and category-canonical sites. A citation-influencing content strategy needs to allocate investment across all three categories with weightings tuned to which engines matter most for the brand's buyer base.

What is source affinity and how is it different from domain authority?

Source affinity is engine-specific — it measures how strongly a single AI engine prefers a specific source domain relative to other engines and relative to the base rate of citations to that domain across the category. Domain authority (in the classic SEO sense) is engine-agnostic — a global score of a domain's link-based authority. A domain can have high domain authority but low source affinity on Claude, or vice versa. In AI search, source affinity is the more predictive metric for how a piece of content will land in a specific engine's answers. Traditional DA is a rough proxy but does not capture the training-data and retrieval-preference nuances that create engine-specific skew.

Can citation trends be tracked over time?

Yes, and they should be. Engine citation trends shift as models are retrained, as retrieval indexes are refreshed, and as new source integrations ship (for example, a new Yelp integration for ChatGPT, or a new Reddit deal for Google). Weekly citation tracking against a stable query set surfaces these shifts as they happen — a source that suddenly gains 5× citation share on Perplexity in a two-week window signals a retrieval-index change worth investigating. Sanbi.ai's citation analytics module tracks per-engine, per-domain citation share weekly and alerts on statistically significant shifts.

How does AI citation tracking help with GEO (Generative Engine Optimization)?

GEO — Generative Engine Optimization — is the discipline of getting content and citations to surface inside AI-generated answers. Citation tracking is the measurement layer that makes GEO iterative rather than blind. Without citation tracking, a GEO program cannot tell whether a new piece of content is being picked up as a source, whether it is being picked up more on one engine than another, or whether it is displacing a competitor's source. Citation tracking turns GEO into a closed loop: publish, measure per-engine citation impact, iterate on the content and distribution channels that moved the needle. See our [GEO playbook](/blog/generative-engine-optimization-geo-playbook-2026-reddit) for the full framework.

What kind of content gets cited most by ChatGPT?

In the analyzed B2B dataset, ChatGPT cited category-canonical manufacturer and OEM domains most heavily — the top three ChatGPT sources accounted for over 30% of its citations combined. This tracks with the wider pattern: ChatGPT with search prefers domain-authoritative, category-canonical sources for factual and technical prompts. Long-form product documentation, technical specification pages, and official reference material outperform blog and social content on ChatGPT specifically. For content investment, that means the highest-leverage move to influence ChatGPT is deep, structured, technically accurate content on your primary domain — not distributed social or third-party placement.

How do I use citation data to prioritize AEO investment?

Use per-engine citation share data to identify the top 10–20 domains cited by each engine for your query set. Then classify each domain as: (a) owned — already yours; (b) earnable — a publication, forum, directory, or platform where you can realistically place content or earn a mention within 3–6 months; (c) unreachable — sources you cannot influence (patents, government, competitor domains). Rank the earnable domains by weighted citation share across the engines you care about. That ranked list is your AEO distribution priority queue for the quarter. Sanbi.ai's citation report exports this list directly.

What is a source-affinity matrix?

A source-affinity matrix is a table (or heatmap) with domains as rows and AI engines as columns. Each cell shows the citation share that engine gives to that domain, color-coded by skew — favored (green, ≥1.5× the average) or avoided (red, ≤0.5× the average), with dashes for 'never cites.' A source-affinity matrix makes engine-specific citation strategy legible at a glance: at a look you can see that (for example) Reddit is a Perplexity play, patents are a Claude play, YouTube is a Perplexity + Gemini play, and manufacturer-official sites are a ChatGPT play. This is the standard output format for citation analytics platforms.

Does Google-owned YouTube get cited more by Gemini than by other engines?

In the analyzed dataset, YouTube was cited more heavily by Perplexity (4.9% share) than by Gemini (1.7% share) for the tracked query set — but Gemini still cited it meaningfully, while ChatGPT and Claude cited it not at all. The wider takeaway is that Google's owner advantage does not automatically translate into Gemini being the dominant YouTube citer for every category — Perplexity's answer format actively embeds YouTube for some query classes in ways Gemini does not. Assume nothing about citation behavior from ownership relationships; measure per query set.

How large a query set do I need to get reliable citation-trend data?

For a stable engine-level citation-share signal within a single category, a query set of 200–500 prompts run weekly against all four major engines is a solid baseline — it typically generates tens of thousands of citations per month, enough for domain-level share to stabilize below noise. For per-domain skew analysis (identifying 1.5×+ affinities), 500–1,000 prompts is more comfortable. Below 100 prompts, citation data is directionally useful but individual per-domain shares will move noisily week to week. Sanbi.ai's tracked prompt sets are configurable at every plan tier — see the [prompt monitoring guide](/blog/prompt-volume-llm-prompt-monitoring-reddit) for how to size a set.

Why is there a big difference between ChatGPT's and Perplexity's total citation counts?

Because engines differ in how many citations they surface per answer. Perplexity is citation-heavy by design — it typically produces 5–15 visible citations per answer. ChatGPT with search enabled surfaces 2–6 citations on average, and often produces answers with no visible citations at all. Claude and Gemini fall in between with their own patterns. In the analyzed dataset, Perplexity produced 39,664 citations from the same query set that generated only 14,721 on ChatGPT — a ~2.7× difference. Citation share (percentages) is the fair comparison across engines, not raw counts.

Where does distributor and marketplace content (e.g., component distributors, aggregators) sit in AI citation trends?

Distributor and aggregator domains show strong engine-specific skew. In the analyzed dataset, distributor / aggregator domains (mouser.com, digikey.com) were heavily cited by Perplexity and Gemini but rarely by Claude. This pattern is common across categories where marketplaces and aggregators sit between OEMs and buyers: the aggregator pages get treated as authoritative comparison sources by live-retrieval engines (Perplexity) and by index-integrated engines (Gemini), while training-data-heavy engines like Claude reach for the OEM source instead. For OEM brands, aggregator presence and listing quality is a meaningful indirect lever on Perplexity + Gemini visibility.

How does Sanbi.ai measure and report engine citation trends?

Sanbi.ai runs each brand's tracked prompt set against ChatGPT, Gemini, Perplexity, and Claude on a scheduled cadence (daily to weekly depending on plan). Every visible citation is parsed, resolved to a root domain, and aggregated into per-engine citation share, cross-engine affinity/skew, and time-series trends. The dashboard surfaces the top cited domains per engine, a domain-×-engine matrix with green/red skew highlighting, and alerts on statistically significant shifts. Reports are exportable as PDF (the format many of the visualizations in this article are drawn from). Full mechanics in the [complete AEO guide](/blog/answer-engine-optimization-aeo-guide-2026).