Back to “Understanding GEO”
Understanding GEO4 min read

ChatGPT, Perplexity, Gemini, Claude: how each picks its sources

Perplexity, ChatGPT, Gemini, Claude: each AI engine has different sourcing habits shaped by its architecture. A careful overview, without overreaching generalizations.

The four major mainstream generative engines don't select their sources the same way, because they rely on different architectures and partnerships. These differences are observable in practice, but they keep evolving: what follows describes trends, not fixed rules.

Why these differences exist

Before going engine by engine, it helps to understand where these gaps come from. Two factors carry particular weight: the share of "internal" knowledge (learned during the model's training) versus real-time web search, and whether or not commercial partnerships exist with content publishers that can influence how often certain sources appear. An engine that relies mostly on internal knowledge reacts slowly to newly published content; an engine that relies mostly on real-time search reacts fast, but then depends heavily on the technical quality of indexing. Another factor worth considering is each model's training knowledge cutoff date: an engine whose training stopped several months before it went into production may miss recent changes in your industry, unless it compensates with active web search. This date varies from one model to the next and shifts with every new version, which makes regular re-checking all the more useful compared to a one-time, frozen assessment.

Perplexity: the conversational search engine

Perplexity is built around real-time web search: each answer draws on a set of search results retrieved at the moment of the query, then synthesized with numbered citations linking back to the source pages. This architecture makes it especially dependent on the freshness and technical accessibility of content (see the lesson on llms.txt and the XML sitemap for the technical side). In practice, Perplexity tends to cite a wide range of sources, including forums, news sites, and documentation pages, making it one of the engines with the highest number of visible sources per answer.

ChatGPT (with browsing): between trained knowledge and live search

ChatGPT combines knowledge learned during the model's training with, when web browsing is enabled, results retrieved live. This duality means your brand's presence in ChatGPT answers depends both on how it was mentioned in the content used for training (past mentions, hard to influence retroactively) and on its current visibility on the web when search is active. OpenAI has content partnerships with several publishers, which can influence how often certain sources appear, without this being exhaustively and publicly documented.

Gemini: integration with the Google ecosystem

Gemini benefits from direct access to Google's search infrastructure, which naturally brings it closer to the logic already at work in classic Google ranking and in AI Overviews (covered in a dedicated lesson on Google AI Overviews). Content that is already well-ranked and technically solid for traditional SEO therefore tends to have more favorable ground to appear in Gemini answers than in other engines built on different logic — without this guaranteeing systematic pickup. Because Gemini answers are surfaced to an enormous default audience through Google Search and the Android ecosystem, even a modest improvement in how it treats your content can reach a disproportionately large number of users compared with a similar improvement on a smaller, opt-in chatbot.

Claude: caution and limited sourcing

Claude, developed by Anthropic, has historically taken a more cautious stance on web browsing and citing external sources than its competitors, tending to rely more on the model's internal knowledge and to state explicit caveats when information can't be verified in real time. This caution can reduce how often it directly cites recent web sources compared to an engine like Perplexity, but it makes it all the more important for your brand to have a presence in widely distributed, widely referenced content that has a better chance of having been folded into the model's knowledge. That includes content picked up by major outlets, indexed documentation, and reference sites likely to be swept into future training data, rather than obscure, low-traffic pages.

What to remember despite the differences

These architectural differences shouldn't lead to a scattered, engine-by-engine strategy: the fundamentals that make content citable (clarity, sourcing, self-contained information, perceived brand authority) hold everywhere. What varies is the relative weight of each factor and the freshness of the sources consulted. A good practice is to regularly observe how your brand appears (or doesn't) in each of these engines' answers on queries representative of your industry, rather than relying on generalities that quickly go stale given how fast these products evolve. Concretely, this means submitting the same panel of questions to Perplexity, ChatGPT, Gemini, and Claude at regular intervals, noting which sources are cited and how your brand is presented within them, then comparing these readings over time rather than relying on a single, isolated observation. This discipline of repeated, engine-by-engine observation is more revealing than a one-off audit: it lets you spot whether a change in your content or in a distribution partnership actually moved the needle on your presence.

Keep in mind too that these behaviors shift with every product update: a snapshot frozen at one point in time ages fast, which is why it's worth following the broader evolution of the research landscape rather than treating any rule about a given engine as final. Treat everything above as a starting map, not a fixed territory: the safest habit is to re-test your own representative queries across all four engines every so often, since a description that felt accurate six months ago may already be out of date.