Back to “Understanding GEO”
Understanding GEO4 min read

What academic research says about GEO

The founding paper on Generative Engine Optimization offers a testing methodology, not a magic formula. What it shows, and its limits.

The most frequently cited reference paper on this topic is titled "GEO: Generative Engine Optimization," published in 2024 by researchers affiliated with Princeton, Georgia Tech, the Allen Institute for AI, and IIT Delhi. It doesn't offer a magic formula for getting cited by an AI, but a methodology for testing and comparing different content optimization techniques against simulated generative engines.

The founding paper: "GEO: Generative Engine Optimization"

This work is generally regarded as one of the first to formalize GEO as a discipline distinct from SEO, by asking the following question: since generative engines synthesize answers from multiple sources rather than simply ranking them, which characteristics of a piece of content increase the likelihood it gets picked up and featured in the generated answer? To answer this, the authors built a set of queries spanning varied domains, meant to serve as a reproducible benchmark for comparing different versions of the same content.

Why this paper mattered for the field

Before this kind of work, discussion around visibility in AI answers stayed largely anecdotal: individual observations, non-reproducible tests, intuitions borrowed from traditional SEO and applied by analogy. The paper's main contribution isn't really any single technique it tested, but the fact of proposing a systematic, shared evaluation framework in the first place, one that let other researchers reproduce, challenge, or refine the results. It's this kind of open methodology, more than any single figure it contains, that became a reference point for subsequent publications on the topic.

How the researchers evaluated generative visibility

The method involves submitting the same query to a generative engine, observing which sources are cited in the produced answer and with what relative weight, then modifying a source's content according to a given technique (adding citations, rephrasing, etc.) before repeating the experiment to see whether its visibility in the generated answer changes. This approach borrows from methods used in information retrieval research to evaluate result relevance, adapted to a case where the final output is no longer a list of links but a synthesized text.

The main families of techniques tested

The researchers studied several categories of content modifications: adding citations and statistics, inserting quotes attributed to recognized figures, simplifying language, enriching text with relevant keywords, and generally improving fluency and structure. The central finding of this research is that some of these techniques have a measurable effect on a source's visibility in the generated answer, but that this effect is neither uniform nor guaranteed: its magnitude varies by the domain tested (health, law, technology...), the engine used, and the nature of the query. In other words, the existence of measurable optimization levers is established, but no technique works universally and consistently.

Since this initial work was published, other research teams and several industry companies have released their own studies or benchmarks on related questions, sometimes with different methodologies and conclusions. This diversity is actually healthy for the field: it signals that no single study, however rigorous, should be treated as the final word on the subject. A serious researcher or practitioner cross-checks several sources against each other rather than relying on a single reference, no matter how often it gets cited.

The limits of this research

This work has important limitations to keep in mind. It reflects the generative engines and models as they existed at the time of publication in 2024: commercial systems evolve rapidly, and results obtained on a test configuration don't necessarily transfer identically to current versions of ChatGPT, Perplexity, Gemini, or Claude. The number of domains and queries tested, while designed to be representative, remains a limited sample compared to the real diversity of use cases. Finally, a controlled research environment differs from a commercial production system, which may incorporate partnerships, safety filters, or business logic that isn't publicly documented and is absent from the experimental setup.

What this means for your content strategy

This kind of academic research should be read as directional confirmation rather than a precise instruction manual: it establishes that optimizing for generative engines isn't a random phenomenon and that measurable levers exist, without providing a formula along the lines of "this technique increases visibility by this percentage" that would hold for your particular case. The right practice is to lean on the broad principles it confirms — the importance of sourcing, covered in the lesson on citations, numbers and evidence, and the importance of a source's perceived credibility, covered in the lesson on E-E-A-T — while directly observing, on your own content and queries, what actually works, rather than relying on a single figure, even one from a serious research paper. As with the market data covered in the lesson on GEO numbers, the best approach remains checking the primary source before repeating a conclusion. Before citing a result from a research paper, make it a habit to read at least the abstract and the methodology section rather than just the title or a press article about it: that's usually where the nuances and limitations live, the ones media coverage tends to smooth over.