ChatGPT, Perplexity and Google AI Overviews cite almost entirely different sources

3 min read Augustyn Głowacki

Three overlapping circles with only a sliver of shared area, standing in for how little AI engines' cited sources overlap.

Any two AI answer engines share only about 9% to 25% of the domains they cite for the same set of queries, and each engine contributes 35% to 68% of domains that nobody else cites (University of Toronto, arXiv 2509.08919, data collected August 2025, measured as pairwise Jaccard similarity by vertical). Being cited by ChatGPT tells you close to nothing about whether Perplexity or Google AI Overviews will cite you for the same question. That's a mechanics problem, not a content problem, and it comes from how differently each engine retrieves.

How different are the citation sets, exactly?

In the Toronto study’s automotive vertical, Claude-GPT pairwise Jaccard was 0.147, Claude-Perplexity 0.251, and GPT-Perplexity just 0.096. In consumer electronics: Claude-GPT 0.150, Claude-Perplexity 0.200, GPT-Perplexity 0.088. Exclusive domain share - the share of cited domains that only one engine ever cites - ran 50.3% to 67.6% depending on engine and category. A separate academic measurement (HKUST-GZ and Rutgers, 55,936 queries, arXiv 2512.09483) corroborates the same shape: only 38% of domains are common between AI engines and traditional search engines, and 37% are unique to AI engines entirely.

Why does ChatGPT overlap least with Google’s rankings?

Because it isn’t retrieving through one index. A researcher extracted ChatGPT’s internal result_source field from its own network payloads and found labels for at least five distinct pipes - a direct SERP call, Bright Data, Oxylabs (both commercial scrapers), a path labeled “labrador”, and Bing - with the mix shifting between accounts and drifting within one account over days (Suganthan Mohanadasan, part 2, 14 July 2026). OpenAI deleted the field from its payloads on 21 July 2026, closing the window that made this observable. The academic anchor for the resulting overlap: ChatGPT’s mean domain overlap with Google was 0.118 and with Bing just 0.041 - actually closer to Google than Bing, the opposite of what a Bing-powered architecture would predict (arXiv 2512.09483).

Which engine tracks classical search rankings most closely?

Perplexity, by a wide margin. Its domain overlap with Google’s top 10 measured at 0.408 in the same academic study, and a separate Semrush analysis of 5,000 keywords put it over 91% at the domain level (Semrush, 21 July 2025) - though that’s a different unit (any-hit rate, not share-of-citations), and the two numbers aren’t directly comparable. What’s consistent across every study I found: Perplexity tracks Google rankings closer than any other major engine, and ChatGPT tracks them least.

What does this mean for a B2B site trying to get cited?

There is no single piece of content that wins everywhere, and optimizing for one engine’s behavior can do nothing for another. On category and comparison questions specifically, the problem compounds: between 82% and 89% of AI citations across engines go to earned media - third-party review sites, trade press, aggregators - not to a brand’s own domain (Muck Rack Generative Pulse, 25 million links, primary study at muckrack.com/blog/what-is-ai-reading-may-2026, distributed via GlobeNewswire). Owned-content work still wins branded and long-tail informational queries. It does not win the category question, and no amount of on-page optimization closes that gap - getting into the third-party sources the engines already read does, and that’s a different project than the one this post is about. Naming which domains each engine actually cites for your own category is a research job, and it is the one that tells you where the coverage has to come from.

Sources

Some of the technology I build with

Newsletter

Get new posts by email

One email when I publish. No drip sequence.