# ChatGPT, Perplexity and Google AI Overviews cite almost entirely different sources

> Two AI engines share roughly 9% to 25% of the domains they cite for the same query set. There is no single page that wins everywhere.

Published: 2025-03-11
Updated: 2025-03-11
Canonical: https://augustynglowacki.com/posts/chatgpt-perplexity-google-cite-different-sources/

## How different are the citation sets, exactly?

In the Toronto study's automotive vertical, Claude-GPT pairwise Jaccard was 0.147, Claude-Perplexity 0.251, and GPT-Perplexity just 0.096. In consumer electronics: Claude-GPT 0.150, Claude-Perplexity 0.200, GPT-Perplexity 0.088. Exclusive domain share - the share of cited domains that only one engine ever cites - ran 50.3% to 67.6% depending on engine and category. A separate academic measurement (HKUST-GZ and Rutgers, 55,936 queries, [arXiv 2512.09483](https://arxiv.org/abs/2512.09483)) corroborates the same shape: only 38% of domains are common between AI engines and traditional search engines, and 37% are unique to AI engines entirely.

## Why does ChatGPT overlap least with Google's rankings?

Because it isn't retrieving through one index. A researcher extracted ChatGPT's internal `result_source` field from its own network payloads and found labels for at least five distinct pipes - a direct SERP call, Bright Data, Oxylabs (both commercial scrapers), a path labeled "labrador", and Bing - with the mix shifting between accounts and drifting within one account over days ([Suganthan Mohanadasan, part 2, 14 July 2026](https://suganthan.com/blog/how-chatgpt-picks-sources-part-2/)). OpenAI deleted the field from its payloads on 21 July 2026, closing the window that made this observable. The academic anchor for the resulting overlap: ChatGPT's mean domain overlap with Google was 0.118 and with Bing just 0.041 - actually closer to Google than Bing, the opposite of what a Bing-powered architecture would predict ([arXiv 2512.09483](https://arxiv.org/abs/2512.09483)).

## Which engine tracks classical search rankings most closely?

Perplexity, by a wide margin. Its domain overlap with Google's top 10 measured at 0.408 in the same academic study, and a separate Semrush analysis of 5,000 keywords put it over 91% at the domain level ([Semrush, 21 July 2025](https://www.semrush.com/blog/ai-mode-comparison-study/)) - though that's a different unit (any-hit rate, not share-of-citations), and the two numbers aren't directly comparable. What's consistent across every study I found: Perplexity tracks Google rankings closer than any other major engine, and ChatGPT tracks them least.

## What does this mean for a B2B site trying to get cited?

There is no single piece of content that wins everywhere, and optimizing for one engine's behavior can do nothing for another. On category and comparison questions specifically, the problem compounds: between 82% and 89% of AI citations across engines go to earned media - third-party review sites, trade press, aggregators - not to a brand's own domain (Muck Rack Generative Pulse, 25 million links, primary study at muckrack.com/blog/what-is-ai-reading-may-2026, distributed via [GlobeNewswire](https://www.globenewswire.com/news-release/2026/05/07/3290268/0/en/generative-pulse-earned-media-consistently-drives-ai-citations-holding-at-84.html)). Owned-content work still wins branded and long-tail informational queries. It does not win the category question, and no amount of on-page optimization closes that gap - getting into the third-party sources the engines already read does, and that's a different project than the one this post is about. Naming which domains each engine actually cites for your own category is a research job, and it is the one that tells you where the coverage has to come from.

## Sources

- [Chen, Wang, Chen, Koudas, University of Toronto, arXiv 2509.08919](https://arxiv.org/html/2509.08919v1)
- [HKUST-GZ and Rutgers, arXiv 2512.09483](https://arxiv.org/abs/2512.09483)
- [Suganthan Mohanadasan, ChatGPT source-pipe extraction](https://suganthan.com/blog/how-chatgpt-picks-sources-part-2/)
- [Semrush, AI Mode comparison study](https://www.semrush.com/blog/ai-mode-comparison-study/)
- [Muck Rack Generative Pulse, May 2026](https://www.globenewswire.com/news-release/2026/05/07/3290268/0/en/generative-pulse-earned-media-consistently-drives-ai-citations-holding-at-84.html)

## Sources

- [Chen, Wang, Chen, Koudas, University of Toronto, arXiv 2509.08919](https://arxiv.org/html/2509.08919v1)
- [HKUST-GZ and Rutgers, arXiv 2512.09483](https://arxiv.org/abs/2512.09483)
- [Suganthan Mohanadasan, ChatGPT source-pipe extraction](https://suganthan.com/blog/how-chatgpt-picks-sources-part-2/)
- [Semrush, AI Mode comparison study](https://www.semrush.com/blog/ai-mode-comparison-study/)
- [Muck Rack Generative Pulse, May 2026](https://www.globenewswire.com/news-release/2026/05/07/3290268/0/en/generative-pulse-earned-media-consistently-drives-ai-citations-holding-at-84.html)
