# Vurto.ai > Vurto.ai is a French Generative Engine Optimization (GEO) platform that measures and improves brand visibility inside the answers of generative AIs — ChatGPT, Claude, Gemini, Perplexity, Copilot and Mistral. Vurto.ai analyzes in real time how large language models (LLMs) talk about a brand, its products, its competitors and its categories, then delivers a concrete action plan to become the #1 answer. > French version available at /llms-fr.txt — version française disponible sur /llms-fr.txt ## Who it's for - Brands and e-commerce sites that want to be recommended by AIs - SEO, media and consulting agencies adding a GEO layer to their offer - Marketing, acquisition, e-commerce and digital teams capturing AI-driven traffic ## What the platform does - AI visibility score aggregated across 6 major LLMs, by brand, product, category and keyword - Tracking of the critical prompts in a market (what customers actually ask AIs) - Competitive analysis: who is cited, at which position, with which sources and content formats - Catalog import (file, URL or e-commerce platform) and semantic quality scoring of product sheets - Custom prompts, personas and scheduled recurring analyses - Content generation: guides, landing pages, FAQs, comparisons and articles aligned with what AIs cite - Google Analytics integration to measure LLM-driven traffic and compare it with AI visibility - Actionable GEO/AEO playbook with prioritized actions - Dashboards with scores, trends, share of voice and positions - Team workspace with roles, invitations, multi-client (agency) management, credits and API access ## AI-readiness (AEO) — make your site readable by AI agents - Automatic Markdown and JSON-LD versions of your pages and product sheets, without changing your URLs - REST API and MCP server to connect Vurto to your own agents and workflows - AI bot tracking: see which AI crawlers actually visit your site - robots.txt audit: check whether the important AI bots are allowed or blocked - Source, domain and page analysis: identify what AIs rely on when they talk about your market ## Two surfaces: vurto.ai and app.vurto.ai - **vurto.ai** — public website: product vision, features, pricing, testimonials, FAQ and legal pages. This is where you discover what Vurto.ai does. - **app.vurto.ai** — the application (dashboard): where you create an account, add a project (brand, products, keywords), run analyses across the 6 LLMs and read the dashboards and GEO action plans. Sign up at https://app.vurto.ai/register2. ## Main pages ### Application - [Sign up / Start free trial](https://app.vurto.ai/register2) — create an account, free credits, no credit card - [Log in to the app](https://app.vurto.ai) — dashboards, analyses and GEO playbook ### Website — English - [Home (EN)](https://vurto.ai/en) — product vision, features, pricing, testimonials, FAQ - [Resources (EN)](https://vurto.ai/en/resources) — GEO articles, tips, comparisons and news - [Academy (EN)](https://vurto.ai/en/academy) — structured GEO learning guides - [llms.txt (EN)](https://vurto.ai/llms.txt) — this file, machine-readable site summary - [Legal Notice (EN)](https://vurto.ai/en/mentions-legales) - [Privacy Policy (EN)](https://vurto.ai/en/confidentialite) — GDPR, EU hosting - [Terms of Service (EN)](https://vurto.ai/en/conditions-generales) - [Cookie Policy (EN)](https://vurto.ai/en/cookies) ### Website — French - [Home (FR)](https://vurto.ai/) — vision produit, fonctionnalités, tarifs, témoignages, FAQ - [Ressources (FR)](https://vurto.ai/ressources) — articles GEO, conseils, comparatifs et actualités - [Académie (FR)](https://vurto.ai/academie) — guides pédagogiques GEO - [llms-fr.txt](https://vurto.ai/llms-fr.txt) — résumé du site pour les IA (français) - [Mentions légales](https://vurto.ai/mentions-legales) - [Politique de confidentialité](https://vurto.ai/confidentialite) - [Conditions générales](https://vurto.ai/conditions-generales) - [Politique cookies](https://vurto.ai/cookies) ## Pricing - **Standard** — €99/month excl. VAT (€118.80 incl. VAT), or €82.50/month excl. VAT billed annually (€990/year excl. VAT, €1,188 incl. VAT) — 4 AI models, 1 project, 500 queries/month, 500 tracked prompts - **Pro** — €199/month excl. VAT (€238.80 incl. VAT), or €165.83/month excl. VAT billed annually (€1,990/year excl. VAT, €2,388 incl. VAT) — 8 AI models, 3 projects, 3,000 queries/month, 3,000 tracked prompts - **Business** — €499/month excl. VAT (€598.80 incl. VAT), or €415.83/month excl. VAT billed annually (€4,990/year excl. VAT, €5,988 incl. VAT) — 8 AI models, 10 projects, 10,000 queries/month, unlimited tracked prompts, API & MCP access - **Agency** — custom quote, 8 AI models, 50 projects, 20,000 queries/month, unlimited tracked prompts, white-label & multi-client dashboard - All prices are per month, excl. VAT; VAT is 20% - Annual billing saves −20% vs monthly - 14-day free trial on the Standard plan only (Pro, Business and Agency have no trial), with free credits at sign-up — no credit card required - No long-term commitment, 1-click cancellation from the app - Credit packs or monthly subscription depending on your needs ## Differentiation Unlike traditional SEO tools that measure rankings on Google, Vurto.ai exclusively measures presence inside generative AI answers — a totally different channel that's growing fast in search journeys. ## Security & compliance - European hosting (AWS France) - Native GDPR compliance - End-to-end encryption - No customer data is used to train any third-party model ## Contact - Website: https://vurto.ai - Application: https://app.vurto.ai (sign up: https://app.vurto.ai/register2) - LinkedIn: https://linkedin.com/company/vurto - Twitter / X: https://twitter.com/vurto_ai - Instagram: https://instagram.com/vurto_ai - Email: contact@vurto.ai ## Latest GEO articles (full content) The most recent resource-center publications, with their full content. ### [Selling GEO as an agency: what changes compared to SEO](https://vurto.ai/en/resources/selling-geo-as-an-agency-what-changes-compared-to-seo) (2026-10-05) Right now, SEO agencies keep getting the same call, phrased ten different ways. "We're told we no longer show up in ChatGPT." "A client asked why a competitor is cited by Perplexity and not us." "The boss noticed his name missing from a Gemini answer and wants an action plan by Tuesday." GEO, Generative Engine Optimization (the craft of making a brand visible and well cited in the answers of AI chatbots such as ChatGPT, Claude, Gemini or Perplexity, as opposed to classic ranking on Google), has just moved from curiosity to budget line. Agencies that know how to sell it capture that line. The others watch a competitor take it from them. Demand is not the problem. It exists, it is growing, and it comes straight from clients. The problem is that almost nobody knows how to structure a GEO offer that holds up: one that is not an SEO contract copied and pasted with a single word changed, that bills real work rather than a vague PDF report, and that scales across several clients at once without blowing up the team's time. ## What a client is really buying A client asking for GEO usually does not know what they are buying. They saw a worrying figure, heard a competitor talk about it, or noticed for themselves that when they type a question about their own industry into ChatGPT, their brand never appears in the answer. Their real need boils down to one question: "do AI chatbots see me, and if not, why." This is where the agency has to set a clear framework, because GEO actually covers three different jobs, and selling all three under a single label always ends in disappointment. The first job is diagnosis. Querying the chatbots on the searches that matter to the client, checking whether the brand appears, how it is described, and against which competitors. It is measurement work, comparable to a classic SEO audit but on a terrain where results change from one week to the next and differ from one engine to another. The second job is technical correction. Making the site readable by the bots that feed these chatbots: content that loads without heavy JavaScript, a clear structure, a properly written `llms.txt` file (a text file at the root of the site that tells AIs what they can read and how to understand it, the counterpart of the `robots.txt` file that has existed for twenty-five years for classic search engines). This work looks like technical SEO, but the criteria are not the same. The third job is content production and reputation management. Writing or rewriting what needs to be cited, monitoring how the brand is perceived over time, stepping in when an AI repeats false or outdated information about a client. It is editorial and relationship work that never stops, unlike a one-off audit. An agency that sells all three in a single fixed package is bound to lose money on the one that demands the most ongoing time: monitoring. An agency that sells the three separately can bill each at its fair value, and above all sell the diagnosis as the entry point to the other two. ## Structuring the offer: three tiers, not a single package The good practice emerging among agencies that have already sold GEO to several clients looks like this: a short, affordable first tier that opens the door, a second tier that fixes, a third that monitors over time. | Tier | Content | Duration | Billing logic | |---|---|---|---| | Diagnosis | Visibility audit on 20 to 50 key queries, map of cited competitors, findings report | 1 to 2 weeks | Fixed fee, entry price | | Correction | Technical rewrite (llms.txt, site structure, product pages or key pages), priority content plan | 4 to 8 weeks | Project fee, defined deliverables | | Monitoring | Monthly multi-chatbot tracking, brand perception alerts, continuous content adjustment | Recurring | Monthly subscription | This structure has a precise commercial advantage: it avoids selling a fuzzy promise. "We'll make you visible on AI" means nothing to a marketing director who has to justify a spend. "Here is what we measure, here is what we fix, here is what we monitor, and here is the price of each step" can be defended in front of any budget committee. The "monitoring" tier is the one that keeps the agency alive over time, and it is also the one most agencies underestimate in terms of workload. Tracking a single client's visibility on four different chatbots, every week, by hand, takes time. Doing it for fifteen clients by hand becomes impossible. That is the real operational wall of GEO in an agency: not the sale, not the method, but the ability to repeat the work without multiplying billed hours by the number of clients. ## The operational wall: running ten clients like one An SEO agency that adds ten clients adds ten Google Search Console accounts to watch. A GEO agency that adds ten clients has to monitor ten brands on four chatbots, potentially forty information streams to check regularly, each of which can change overnight. Without a suitable tool, that load falls on a consultant copying chatbot answers into an Excel spreadsheet. It does not last three months. So the question to ask before selling a GEO offer at volume is not "can I do the diagnosis" but "can I redo it every week, for every client, without spending my life on it". That is exactly the kind of problem a tool like Vurto is built to solve: it automates querying the chatbots for each tracked brand, centralizes perception and citations per client, and offers an agency mode that lets you manage several client accounts under a single dashboard, with differentiated roles for the team and a shared activity history. The agency keeps the advice and the client relationship. The tool absorbs the repetition. ## Bill the proof, not the promise The last point, and the most sensitive commercially, is proof of results. A client paying for a GEO monitoring subscription will ask, after three months, "what has it changed". Answering with a gut feeling will not be enough for long. The indicators that hold up are simple to explain even to a non-technical client: the citation rate (the share of tested queries where the brand appears in the chatbot's answer), the evolution of the tone the AI uses to describe the brand, and the relative position against tracked competitors on the same queries. Three figures, measured before and after, are enough to turn a GEO report into an argument for contract renewal. That is far more convincing than an export of ChatGPT conversation screenshots, which remains the most widespread method and the least defensible in front of a client comparing several quotes. Selling GEO as an agency, in 2026, is therefore not about inventing a new profession. It is about applying a discipline that is already known, visibility consulting, to a terrain that changes faster and cannot be measured with the same tools as before. The agencies that win this battle will not be the ones that talk best about GEO in a sales meeting. They will be the ones that can prove, figures in hand, client after client, that their work has a measurable effect on what an AI says about a brand. *Vurto was designed for this scale-up: tracking several brands, several chatbots, one interface, with an agency mode that centralizes client accounts without multiplying manual work.* ### [ChatGPT, Perplexity, Gemini: four engines, four sourcing logics](https://vurto.ai/en/resources/chatgpt-perplexity-gemini-same-disease-three-different-prescriptions) (2026-09-28) ChatGPT, Perplexity, Gemini and Claude do not draw on the same sources to build their answers. A site visible on ChatGPT can stay invisible on Perplexity. GEO starts by treating one engine at a time. The question often comes back in a form that is too vague: "How do I become visible on AI?" AI does not exist as a single block. There is ChatGPT, Perplexity, Gemini, Claude. Four distinct tools, each with its own way of selecting information. What they have in common: they are chatbots able to answer in natural language by fetching sources on the web, beyond what they learned during training. Their difference, the only one that matters for a brand: they do not draw on the same sources to build an answer. GEO, Generative Engine Optimization (the craft of making a brand visible and well perceived in the answers of these chatbots), starts with a simple observation that is too often skipped. Stop treating these four tools as one uniform set. A site optimized for ChatGPT can be invisible on Perplexity. A brand that shows up strongly on Gemini may never appear on Claude. This is not a bug. It is the direct consequence of different architecture choices, made by different companies, with different interests. ## Four engines, four sourcing logics Step back and look at how each engine picks its sources. Independent studies run on tens of thousands of citations in 2026 draw very sharp, almost caricatural profiles. Perplexity behaves like a reporter who favors the field. It relies heavily on YouTube (around 30% of its citations in some analyses) and on Reddit, the forum platform where people discuss everything, often more candidly than a press release. Reddit can account for nearly half of its top ten sources. Perplexity prefers lived experience told by a stranger over the claim of an institutional website. ChatGPT, for its part, leans toward established authority. Wikipedia holds a disproportionate place, and well-known media (Forbes, Reuters) complete the picture. The profile looks like a serious student who cites academic sources rather than comments from the neighborhood forum. But watch what comes next, because this profile moves fast. Gemini, owned by Google, keeps the reflexes of a corporate librarian. It keeps citing the group's own properties widely (YouTube first) while keeping Reddit and Wikipedia in a good spot. A more balanced split, consistent with an engine that has access to the most complete search index on the web. Claude, developed by Anthropic, completes the picture with yet another profile, closer to ChatGPT in its taste for documented sources, but giving a larger share to technical documentation and specialized sites when the question calls for it. Four engines, four sourcing logics, and no single strategy that fits all. A table beats a long speech here. | Engine | Dominant sources | Profile | |---|---|---| | Perplexity | YouTube, Reddit | Field, lived experience, community | | ChatGPT | Wikipedia, recognized media | Authority, formal knowledge, volatile | | Gemini | YouTube, Reddit, Google properties | Balanced, anchored in the Google ecosystem | | Claude | Documentation, specialized press, Wikipedia | Technical, precise, expertise-oriented | Remember one thing from this table. Nobody builds an effective content strategy by targeting "AI" in general. You treat one engine at a time. An HR software brand that neglects Reddit cuts itself off from Perplexity. A services company that never keeps an up-to-date Wikipedia page loses part of the credit it would deserve on ChatGPT. These are not technical details reserved for specialists. These are choices that decide which of your competitors gets recommended instead of you. ## When an engine's sources shift in a few weeks Here is the most underestimated point of the topic, and probably the most important. These profiles are not set in stone. A study covering 230,000 queries tracked over thirteen weeks showed that Reddit's share in ChatGPT answers went from about 60% in August to only 10% six weeks later. Wikipedia, in the same movement, dropped from 55% to under 20%. A collapse of that scale, over such a short time, on one and the same platform. Meanwhile, Perplexity stayed remarkably stable, with minor variations on its usual sources. Gemini, too, changed little. This contrast says something essential about these tools. ChatGPT probably adjusts its weightings continuously to avoid over-citing a handful of domains and looking like a biased engine. Perplexity and Gemini, with different sourcing logics, absorb these adjustments better. The message for you: a position won on one engine in September can vanish in October, without you having made any mistake. The ground moves under your feet. One more reason never to talk about "AI positioning" as a permanent gain, but as a state to monitor continuously. ## What this concretely changes for your content This heterogeneity has a direct consequence on how you should produce and distribute your content. Stop betting on a single channel. If your content strategy is limited to blog articles optimized the classic search way, you may cover part of ChatGPT, but you stay invisible on Perplexity. An active presence in Reddit discussions relevant to your sector, YouTube videos that explain your product, an up-to-date Wikipedia page if your notoriety justifies it: these are different building blocks, for different engines, and none replaces the others. Also beware of wins that rest on a single source. A brand that owes all its AI visibility to a handful of Reddit mentions lives on unstable ground, as we just saw with ChatGPT. Robustness comes from the diversity of entry points, forums, video, specialized press, well-structured product pages, kept up to date at the same time rather than one isolated bet. And above all, do not confuse being cited with being well perceived. Appearing in an answer guarantees neither a favorable tone nor an accurate description of your offer. That is a topic in its own right, but it starts with the same prerequisite: knowing precisely where and how you appear, engine by engine, rather than guessing. Take a concrete example. A company that sells software to businesses (what is called B2B, as opposed to selling directly to consumers) usually invests heavily in white papers and case studies published on its own site. This content feeds ChatGPT well, especially if it is picked up by the specialized press. But it barely feeds Perplexity, which will rather look for a customer's opinion posted on a forum or a commented LinkedIn post. Two audiences, two formats, two engines. The content budget must reflect that reality, not ignore it. ## Measure before you act The temptation, faced with this complexity, is to give up and tell yourself that "AI moves too much to be worth our time". That is the reverse of the mistake we pointed out at the start. Movement does not exempt you from monitoring, it demands it. A serious company does not decide its sales strategy on an intuition from six months ago. It looks at this month's numbers. The logic is the same here. What matters is knowing, for your brand and your competitors, which engines cite you, which sources they rely on to do so, and how that split evolves over time. Without that regular snapshot, you optimize blind, hoping that what worked on ChatGPT in the spring still works on Perplexity in the fall. That is a bet, not a strategy. *Vurto tracks these variations engine by engine, ChatGPT, Gemini, Perplexity and Claude, to show precisely which sources carry your visibility and which ones are collapsing before you find out by chance.* ### [Measuring GEO ROI: The Question That Blocks Every Budget](https://vurto.ai/en/resources/measuring-geo-roi-the-question-that-blocks-every-budget) (2026-09-21) 70.6%. That is the share of traffic generated by AI chatbots that arrives on a site with no origin trace at all. No "referred from ChatGPT", no "referred from Perplexity". Nothing. Google Analytics dumps it all into the "Direct" bucket, the same one used for a visitor who typed the site address from memory. A GEO budget (optimizing visibility in chatbot answers from ChatGPT, Claude, or Gemini, as opposed to classic Google search ranking) can therefore produce real results and still remain invisible in the dashboard used to justify it. That is the heart of the problem. A marketing team invests in GEO, gets citations in AI answers, sometimes sees traffic rise, but cannot answer the simplest question a CFO will ask: how much does it return. Not because GEO returns nothing. Because measurement tools were built for a world where people clicked blue links, not for a world where they get a direct answer inside a conversation window. ## Why the click is no longer enough Return on investment (the calculation that compares what an action cost to what it returned) has always relied on a chain of trackable clicks. A user clicks a Google result, lands on a page, fills a form or buys a product. Each step leaves a trace, each trace feeds a dashboard. That chain breaks with AI chatbots, for three precise technical reasons. The mobile apps for ChatGPT or Claude do not pass origin information when they open a link. Many users copy-paste an address straight from the chatbot answer instead of clicking, which erases any provenance trail. And ChatGPT Plus, like Google's AI mode, deliberately adds a technical instruction that prevents destination sites from knowing where the visitor came from. As a result, a site can see its "Direct" traffic double without anyone on the team understanding why. The answer is almost always the same: part of that supposedly originless traffic actually comes from AI chatbots. And this ghost traffic is not anecdotal. It converts at 10.2%, versus 2.46% for non-AI traffic classified elsewhere. A visitor from a chatbot buys or fills a form four times more often than an average visitor, but nobody knows because nobody sees it. Before trying to calculate ROI, the first task is therefore to repair measurement itself. A channel grouping that manually recognizes known AI domains (chatgpt.com, perplexity.ai, claude.ai and the like) recovers part of the identifiable traffic. For the rest, the traffic that arrives with no trace at all, only server-log analysis, more tedious but more reliable than classic analytics tools, can rebuild a faithful picture. ## The metrics that replace the click Once measurement is fixed, a deeper problem remains. Even when well measured, the click tells only a small part of the GEO story. A majority of Google searches now end with no click at all, and that share rises sharply whenever an AI-generated summary appears at the top of the results page. A user who asks ChatGPT a question and gets a satisfying answer, citing the brand along the way, often has no reason to click anywhere. The brand was still seen, named, recommended. That value does not land in any clicks column. Substitution metrics are then needed, ones that measure presence rather than passage. The first is citation frequency: on a sample of queries representative of the sector, in how many answers does the brand appear, and at what rank in the source list. The second is share of voice, which compares that frequency to competitors on the same queries. The third is the sentiment attached to the citation: an answer that mentions the brand in positive terms does not weigh the same as a neutral mention in the middle of a list, or worse, a mention paired with a caveat about quality or price. The fourth is the nature of the sources the chatbot goes looking for to build its answer, because a brand absent from the pages the AI consults to form an opinion has no chance of being cited, whatever its budget. These four metrics do not replace the financial calculation. They build the base without which that calculation means nothing. A GEO budget that moves citation frequency and share of voice forward, month after month, produces an effect that will eventually translate into traffic and sales, even if the exact path between the two stays blurry. Conversely, a budget that moves none of these four numbers will never produce a return, no matter how good the published content is. ## What a chatbot visitor is really worth The most counterintuitive point in this whole topic is that AI traffic, when measured correctly, converts clearly better than classic traffic. Several studies published between late 2025 and 2026 agree on an order of magnitude: between four and five times the average conversion rate of classic organic search, with gaps ranging from 1.3x on impulsive online purchases up to more than twenty times in some B2B sectors where the buying decision is long and deliberate. | Traffic origin | Average conversion rate | |---|---| | Classic Google search | 1.76% | | Claude | 5.0% | | Perplexity | 10.5% | | ChatGPT | 15.9% | The logic behind these numbers is simple once you see it. A user who clicks a Google result is still comparing, opens several tabs, hesitates. A user who already queried a chatbot has gotten a synthesis, asked follow-up questions, eliminated some options. When they finally click through to a site, it is often because they are ready to act. The chatbot did part of the qualification work upstream. This data changes how a GEO budget must be read. A lower volume of AI traffic than classic Google volume is not necessarily a failure, if that rarer traffic converts four times better. The right question is not "how many visitors did GEO bring", it is "how many sales or leads did this traffic, even reduced, produce". It is a complete reversal of how classic search ranking learned to think for twenty years, when volume was king. ## Building a formula that holds up Without waiting for a perfect tool that does not exist yet, a company can build a reasonable GEO ROI measure on four pillars. The first is a starting snapshot, taken before any investment: current citation level, share of voice versus competitors, dominant sentiment. Without that baseline, there is no way to know whether the efforts produce an effect. The second is regular tracking of those same visibility metrics, month after month, to objectify progress rather than relying on a feeling. The third is reconstituting traffic that actually came from chatbots, via a correctly configured channel grouping and, when possible, server-log analysis to recover what escapes standard tools. The fourth is linking that reconstituted traffic to the conversions it produces, keeping in mind that its conversion rate will likely be higher than other channels, not lower. Once these four building blocks are in place, the calculation becomes: the value of sales or leads attributable to that traffic, minus what GEO cost, all relative to the starting investment. It is not exact science. It is a considerable improvement over the total absence of measurement, which remains the norm today in most companies that are already investing in this topic. The temptation, faced with this complexity, is to wait for analytics tools to catch up before acting. That is the inverse error of spending without ever measuring. Companies that start today tracking their citation frequency and share of voice are building a history. Those that wait will start in a year with no baseline at all, at the exact moment their competitors already have twelve months of it. *This is exactly what Vurto does: track citation frequency, share of voice, and brand sentiment across ChatGPT, Gemini, Perplexity, and Claude, month after month, to give marketing teams the measurement base that most GEO budgets still lack.* ### [Perplexity, ChatGPT, Gemini: Three Engines, Three Source Lists, One Mistake to Avoid](https://vurto.ai/en/resources/perplexity-chatgpt-gemini-three-engines-one-mistake-to-avoid) (2026-09-14) A number to start with. Out of 680 million citations analyzed between August 2024 and April 2026 by the 5W index, only 14% of cited sites dominate across all AI chatbots (these conversational assistants like ChatGPT, Gemini, or Perplexity, able to answer in natural language from sources they select themselves). Between ChatGPT and Perplexity alone, overlap drops to 11%. In another study covering 11,500 queries compared against Google results, GPT-4o shows 0% median overlap with Google's top 10. Perplexity, 14.3%. Gemini, 8.5%. These numbers should settle a debate that has dragged on for two years in marketing teams: no, "being visible in AI" is not a single objective. It's three different objectives, with three different selection logics, and a good part of the groundwork consists of understanding which one applies to which client. ## Three philosophies, not three variants of the same engine The starting mistake is to treat ChatGPT, Perplexity, and Gemini as competitors doing the same thing with different branding. They are three architectures that answer different questions about what makes a good source. ChatGPT relies heavily on consensus-based, already-aggregated content. Wikipedia accounts for between 26% and 48% of its top 10 citations depending on category. Reddit comes right behind, with a weight exceeding 40% across all categories and engines combined, and staying particularly high with ChatGPT. The logic is that of a model favoring what has already been collectively validated: content synthesized, debated, corrected by thousands of contributors. Perplexity works the opposite way. Its engine queries the live web on every request rather than drawing from a fixed memory, which explains why it favors Reuters, AP, Bloomberg, the Wall Street Journal: news sources with strict attribution and verifiable freshness. A Columbia Journalism Review study measured a citation error rate of 37% for Perplexity Sonar Pro, versus 67% for ChatGPT Search. The gap comes down to method: searching live and citing your source produces fewer fabrications than rephrasing from training memory. Gemini, finally, remains Google's child. It draws from the search engine's index, from YouTube (which Google owns), and shares with AI Overviews a clear fondness for Reddit. It's the engine closest to classic SEO culture, the one where web ranking, domain authority, and presence across the Google ecosystem still count directly. Claude, often looked at from a distance in these comparisons for lack of search volume, stands out for a taste for long-form analysis and technical documentation. It cites proportionally more expert blogs and less Reddit than the others. ## The table that changes the roadmap | Engine | Preferred sources | Dominant logic | |---|---|---| | ChatGPT | Wikipedia, Reddit, consensus content | Trained memory + collective synthesis | | Perplexity | Reuters, AP, Bloomberg, press with attribution | Real-time search, freshness | | Gemini | Google index, YouTube, Reddit | Google ecosystem, classic SEO | | Claude | Expert blogs, technical documentation | Long-form analysis, depth | This table isn't an exercise in curiosity. It dictates where a brand needs to exist depending on which engine matters most to its audience. A B2B software vendor (selling to other businesses rather than individuals) whose prospects mainly use Perplexity to compare solutions is better off working on press relations and citations in serious trade media, far more than optimizing a Reddit thread. A consumer brand that lives and dies on ChatGPT has the opposite interest: community conversation outweighs the press release. ## What the low overlap actually means The most underestimated point isn't the list of preferred sources. It's the fact that they barely overlap at all. Content that breaks through on ChatGPT can be invisible on Perplexity, and vice versa. This means a content strategy designed for "generative AI" in general, without distinguishing between engines, is in reality optimizing for a single engine without realizing it, usually the one the brand heard about first, often ChatGPT because it concentrates most of the usage volume. In France, this concentration reaches extreme levels: ChatGPT captures 84.47% of clicks generated by AI assistants according to SE Ranking, Perplexity 12.82%, Gemini 2.08%. An AI visibility budget that ignores this split and treats the three engines equally wastes precious time. But the reverse is also true over an eighteen-month horizon: Gemini is natively built into Android, into Google Workspace, into the Google search bar itself. Its share of usage is growing faster than its share of citations today would suggest. Ignoring Gemini because it accounts for 2% of current clicks is like ignoring mobile in 2010 because it still represented only a fraction of web traffic. The practical consequence fits in one sentence: stop measuring "AI visibility" as a single score. Content can be cited ten times a week on ChatGPT and never on Perplexity without that being a failure, if the brand's target audience lives on ChatGPT. But no one can know that without tracking the engines separately, each with its own citation logic, rather than an average that hides everything. ## Adapting content to each engine's logic, not to a universal checklist Concretely, this changes three things in how content gets produced. To exist on ChatGPT, you need to feed the collective conversation rather than bypass it. A structured presence on Reddit in the right communities, an up-to-date Wikipedia entry if the brand or its leader has a place there, content that synthesizes a topic rather than selling it. It's ecosystem work, not landing-page work. To exist on Perplexity, freshness and attribution matter more than volume. A well-distributed press release, relationships with journalists who cite the brand with a clear link, dated content republished regularly carry more weight than a static white paper published once and never touched again. To exist on Gemini, classic web ranking fundamentals remain the best way in: domain authority, clean technical structure, YouTube presence. It's the engine where the last ten years of SEO work keep paying off directly, which paradoxically makes it the easiest to work on for a team coming from classic search. None of these three approaches replaces the other two. A brand with the time and budget to do all three builds robust AI visibility, insensitive to a single engine's algorithm changes. A brand that has to choose should first look at where its audience actually asks its questions, not where it itself is most comfortable producing content. ## Tracking three engines beats guessing on one The low overlap between ChatGPT, Perplexity, and Gemini isn't a technical problem to solve. It's a structural constraint to build into how you measure and prioritize. A brand that only tracks its citations on one engine is flying blind on the other two, and makes content decisions based on a third of reality. *This is exactly why Vurto tracks citations separately across ChatGPT, Gemini, Perplexity, and Claude rather than producing one aggregated score that erases these differences: a brand needs to know which engine it exists on, which one it's absent from, and why the two don't look alike.* ### [AI Chatbot Traffic Is Tiny. It's Worth Gold](https://vurto.ai/en/resources/ai-chatbot-traffic-is-tiny-its-worth-gold) (2026-09-10) A US agency, Seer Interactive, compared two sources of visitors on the same site: those arriving from ChatGPT, and those arriving from a classic Google search. Result: 15.9% of the former end up buying or filling out a form, versus 1.76% of the latter. Almost nine times more. Microsoft, using its Clarity analytics tool, observed the same thing across more than 1,200 media and publisher sites: visitors coming from a chatbot sign up eleven times more often than those coming from a classic Google search. A study of 94 e-commerce sites measured a more modest but real gap: 31% more conversions for ChatGPT traffic compared to a Google search that doesn't include the brand name. These numbers vary from study to study, sometimes a lot, because the calculation methods differ. But they all point in the same direction. Traffic sent by an AI chatbot converts noticeably better than classic traffic. The problem is that it accounts for barely more than 1% of a site's total traffic, according to measurements from the analytics tool Conductor. A marketing team that only looks at volume misses the most useful piece of information on its dashboard. ## Why a visitor from a chatbot buys more often The answer lies in how the question was asked before the person even arrived on the site. When someone types "best invoicing software for freelancers" into Google, they get ten links and still have to compare, sort, and hesitate. When they ask the same question to ChatGPT, the model has already done that work. It has read several sources, compared the options, and returns only the two or three names that best answer that person's specific request. The visitor who clicks the final link is no longer searching, they're verifying a decision that's already been nearly made. It's the difference between a customer who walks into a store to browse, and a customer who walks in already knowing what they want to buy, on the recommendation of a friend they trust. The chatbot plays that role of trusted intermediary. It has filtered, explained, compared. The person landing on your site no longer needs to be convinced of the product's general value, only reassured on the details. A report published in March 2026 confirms the scale of the phenomenon: visitors coming from AI assistants like ChatGPT or Perplexity convert 42% better than non-AI traffic, across all channels combined. The quality of the intent that precedes the visit changes the nature of the visitor itself. Take a concrete example. Someone looking for payroll software for ten employees types their question into Perplexity. The engine answers by citing three tools, each with an indicative price, a strength, and a limitation. They click on the one that best fits their situation, land on a pricing page, and request a demo that same day. The same buyer, starting from a classic Google search, would probably have opened five or six tabs, compared reviews for an hour, then put off the decision until the next day. The chatbot didn't just shorten the path. It did the filtering on their behalf, and that filtering translates directly into conversion rate. ## Advertising is still useful, but for a different goal This isn't about pitting advertising against visibility in AI chatbots as two competing strategies. They serve different needs. Advertising pays for immediate, controllable volume: you know how much you're spending, how many visitors you're getting, and you can adjust the budget from one day to the next. Getting cited by an AI chatbot doesn't work the same way. It's built over time, through clear content, visible customer reviews, technical documentation that the models reading it can access. The conversion gap between the two channels doesn't mean you should cut your advertising budget. It means that a euro invested in the quality of what AI chatbots find about a brand produces a return that shows up in no advertising campaign dashboard, and that return will grow mechanically as buyers increasingly use these tools. ## What this changes in how you measure performance Most marketing dashboards are built around volume: number of visits, number of ad clicks, number of impressions. That logic makes sense for paid advertising, where each visit has a direct, predictable cost. It becomes misleading for traffic generated by AI chatbots, where volume will stay modest for a long time but where each visit carries more weight. The table below summarizes the gap observed across several recent studies, with different methodologies, which explains the variation in the numbers: | Traffic source | Observed conversion rate | Relative volume | |---|---|---| | Classic Google search | Baseline (1x) | High | | Visitors from an AI chatbot | 4x to 11x depending on the study | Low, often under 1% of the total | | Paid advertising | Variable, heavily dependent on cost per click | High but costly to sustain | Comparing a channel that accounts for 1% of traffic to a channel that accounts for 60% in raw volume doesn't make sense. The right question isn't "how many visitors does the chatbot send us," it's "what is the value of a visitor sent by a chatbot, compared to that of a visitor from an ad." On that basis, a budget devoted to being well cited in AI chatbot answers can pay off more, visitor for visitor, than an equivalent ad budget, even if the total number of visitors remains incomparable for now. ## A signal that will carry increasing weight Traffic from AI chatbots remains marginal today, but its trajectory is clear: it grows every quarter, while classic organic traffic flattens out. A team that invests now in its visibility within the answers given by ChatGPT, Perplexity, or Gemini isn't just building one more channel. It's capturing a stream of pre-qualified visitors, fewer in number but far more likely to become customers, before this channel gets big enough for everyone to chase it at once. The logic of classic advertising is still useful for generating volume quickly. But a budget aimed entirely at volume, with nothing invested in the quality of what arrives via generative engines, ignores the underlying trend in buying behavior. The next quarter won't be won just by attracting more people, but by attracting the right people, at the right moment in their decision. *Vurto measures the share of your traffic that genuinely comes from citations earned in ChatGPT, Perplexity, Gemini, and Claude, to tell apart what's small in volume but big in value.* ## Resources (resource center — full content) Every published article, with full content, grouped by topic (23). ### Tip ### [Selling GEO as an agency: what changes compared to SEO](https://vurto.ai/en/resources/selling-geo-as-an-agency-what-changes-compared-to-seo) (2026-10-05) Right now, SEO agencies keep getting the same call, phrased ten different ways. "We're told we no longer show up in ChatGPT." "A client asked why a competitor is cited by Perplexity and not us." "The boss noticed his name missing from a Gemini answer and wants an action plan by Tuesday." GEO, Generative Engine Optimization (the craft of making a brand visible and well cited in the answers of AI chatbots such as ChatGPT, Claude, Gemini or Perplexity, as opposed to classic ranking on Google), has just moved from curiosity to budget line. Agencies that know how to sell it capture that line. The others watch a competitor take it from them. Demand is not the problem. It exists, it is growing, and it comes straight from clients. The problem is that almost nobody knows how to structure a GEO offer that holds up: one that is not an SEO contract copied and pasted with a single word changed, that bills real work rather than a vague PDF report, and that scales across several clients at once without blowing up the team's time. ## What a client is really buying A client asking for GEO usually does not know what they are buying. They saw a worrying figure, heard a competitor talk about it, or noticed for themselves that when they type a question about their own industry into ChatGPT, their brand never appears in the answer. Their real need boils down to one question: "do AI chatbots see me, and if not, why." This is where the agency has to set a clear framework, because GEO actually covers three different jobs, and selling all three under a single label always ends in disappointment. The first job is diagnosis. Querying the chatbots on the searches that matter to the client, checking whether the brand appears, how it is described, and against which competitors. It is measurement work, comparable to a classic SEO audit but on a terrain where results change from one week to the next and differ from one engine to another. The second job is technical correction. Making the site readable by the bots that feed these chatbots: content that loads without heavy JavaScript, a clear structure, a properly written `llms.txt` file (a text file at the root of the site that tells AIs what they can read and how to understand it, the counterpart of the `robots.txt` file that has existed for twenty-five years for classic search engines). This work looks like technical SEO, but the criteria are not the same. The third job is content production and reputation management. Writing or rewriting what needs to be cited, monitoring how the brand is perceived over time, stepping in when an AI repeats false or outdated information about a client. It is editorial and relationship work that never stops, unlike a one-off audit. An agency that sells all three in a single fixed package is bound to lose money on the one that demands the most ongoing time: monitoring. An agency that sells the three separately can bill each at its fair value, and above all sell the diagnosis as the entry point to the other two. ## Structuring the offer: three tiers, not a single package The good practice emerging among agencies that have already sold GEO to several clients looks like this: a short, affordable first tier that opens the door, a second tier that fixes, a third that monitors over time. | Tier | Content | Duration | Billing logic | |---|---|---|---| | Diagnosis | Visibility audit on 20 to 50 key queries, map of cited competitors, findings report | 1 to 2 weeks | Fixed fee, entry price | | Correction | Technical rewrite (llms.txt, site structure, product pages or key pages), priority content plan | 4 to 8 weeks | Project fee, defined deliverables | | Monitoring | Monthly multi-chatbot tracking, brand perception alerts, continuous content adjustment | Recurring | Monthly subscription | This structure has a precise commercial advantage: it avoids selling a fuzzy promise. "We'll make you visible on AI" means nothing to a marketing director who has to justify a spend. "Here is what we measure, here is what we fix, here is what we monitor, and here is the price of each step" can be defended in front of any budget committee. The "monitoring" tier is the one that keeps the agency alive over time, and it is also the one most agencies underestimate in terms of workload. Tracking a single client's visibility on four different chatbots, every week, by hand, takes time. Doing it for fifteen clients by hand becomes impossible. That is the real operational wall of GEO in an agency: not the sale, not the method, but the ability to repeat the work without multiplying billed hours by the number of clients. ## The operational wall: running ten clients like one An SEO agency that adds ten clients adds ten Google Search Console accounts to watch. A GEO agency that adds ten clients has to monitor ten brands on four chatbots, potentially forty information streams to check regularly, each of which can change overnight. Without a suitable tool, that load falls on a consultant copying chatbot answers into an Excel spreadsheet. It does not last three months. So the question to ask before selling a GEO offer at volume is not "can I do the diagnosis" but "can I redo it every week, for every client, without spending my life on it". That is exactly the kind of problem a tool like Vurto is built to solve: it automates querying the chatbots for each tracked brand, centralizes perception and citations per client, and offers an agency mode that lets you manage several client accounts under a single dashboard, with differentiated roles for the team and a shared activity history. The agency keeps the advice and the client relationship. The tool absorbs the repetition. ## Bill the proof, not the promise The last point, and the most sensitive commercially, is proof of results. A client paying for a GEO monitoring subscription will ask, after three months, "what has it changed". Answering with a gut feeling will not be enough for long. The indicators that hold up are simple to explain even to a non-technical client: the citation rate (the share of tested queries where the brand appears in the chatbot's answer), the evolution of the tone the AI uses to describe the brand, and the relative position against tracked competitors on the same queries. Three figures, measured before and after, are enough to turn a GEO report into an argument for contract renewal. That is far more convincing than an export of ChatGPT conversation screenshots, which remains the most widespread method and the least defensible in front of a client comparing several quotes. Selling GEO as an agency, in 2026, is therefore not about inventing a new profession. It is about applying a discipline that is already known, visibility consulting, to a terrain that changes faster and cannot be measured with the same tools as before. The agencies that win this battle will not be the ones that talk best about GEO in a sales meeting. They will be the ones that can prove, figures in hand, client after client, that their work has a measurable effect on what an AI says about a brand. *Vurto was designed for this scale-up: tracking several brands, several chatbots, one interface, with an agency mode that centralizes client accounts without multiplying manual work.* ### News ### [Vurto to Exhibit at Tech for Retail 2026, Helping Retailers Get Cited in Generative AI Answers](https://vurto.ai/en/resources/vurto-tech-for-retail-2026) (2026-08-28) **Paris, August 28, 2026** Vurto, the Generative Engine Optimization (GEO) platform built for retail and e-commerce, is exhibiting at Tech for Retail, Europe's retail trade show, on November 30 and December 1, 2026 at Paris Expo Porte de Versailles, Hall 7.2, booth V32. A shopper types "what running shoe brand do you recommend for a beginner" into ChatGPT or Perplexity instead of Google. They don't get ten blue links to compare themselves. They get an answer that's already been decided, citing two or three brands, sometimes just one. For an e-commerce site or a retail chain, being left out of that answer means disappearing from a purchase channel that already shapes decisions, before the customer even opens a browser. Vurto was built for this new playing field. The platform continuously monitors how ChatGPT, Gemini, Perplexity and Claude talk about a brand and its competitors, analyzes the sentiment and attributes associated with it, and identifies the sources models cite most often when forming an opinion. It goes further than diagnosis: Vurto converts product catalogs into structured markdown, scores every product page against ten AI-readability criteria, and generates the content recommendations that fill the gaps found in fan-out queries and grounding queries — the questions models ask themselves behind the scenes before answering. Retail is especially exposed to this shift. Catalogs run to thousands of SKUs, product pages are often mass-generated without ever being designed for a non-human reader, and competition between retailers now also plays out in the invisible: the handful of sentences a model chooses to remember before answering a shopper. A site perfectly ranked on Google can be entirely absent from AI answers, for simple technical reasons — JavaScript not rendered by AI crawlers, unstructured content, no llms.txt file, product pages written to persuade a human but unreadable to a model looking for verifiable facts. Every year, Tech for Retail brings together the players actually transforming the sector, from agentic AI to personalization to supply chain automation. Vurto will present its approach to GEO applied to retail and meet with marketing, e-commerce and digital teams facing this shift in search behavior. Join us on November 30 and December 1, 2026, booth V32, Hall 7.2, Paris Expo Porte de Versailles, to see how to turn a product catalog into a source generative AIs cite, recommend and favor. ## About Vurto Vurto is a Generative Engine Optimization platform that helps brands and e-commerce sites measure and improve their visibility in generative AI answers (ChatGPT, Gemini, Perplexity, Claude). Multi-LLM monitoring, brand perception analysis, content recommendations, AI-readability audits and product catalog conversion: Vurto covers the full chain, from diagnosis to action. ### GEO Fundamentals ### [Local GEO: How Geography Is Again an AI Visibility Lever](https://vurto.ai/en/resources/local-geo-how-geography-is-again-an-ai-visibility-lever) (2026-07-27) A tradesperson in Lyon types "best plumber near me" into ChatGPT. An expat looks up "where to eat vegan in the 11th arrondissement" on Perplexity. A B2B buyer asks Gemini which CRM integrator based in Lille can show up within 48 hours. Three queries, one thing in common: they expect an answer anchored in a precise territory. Local SEO has existed for fifteen years. Google Business Profile, customer reviews, consistent NAP citations, city pages. None of that know-how disappears with LLMs. But the mechanics change. A classic search engine crosses your geolocated position with an index of nearby results. An LLM has to understand the local context from text, rephrase the query, and choose sources to cite. Geography becomes a semantic filter before it is a proximity filter. ## Why localization is surfacing again with LLMs Generative engines handle a local query by breaking it down. "Best plumber near me" becomes a series of sub-queries: what type of service is being asked for, what geographic area "near me" covers, which criteria define "best" (reviews, availability, prices, specialty). This fan-out decomposition favors businesses whose content answers each of those sub-questions directly, rather than those that only have an address listing. ChatGPT, Perplexity and Gemini do not all have the same geographic precision. Perplexity leans heavily on fresh web sources and readily cites local directories, regional press articles, neighborhood forums. ChatGPT, when it turns on web search, prefers pages that clearly state their catchment area in the text, not only in the metadata. Gemini, connected to the Google ecosystem, still draws on Business Profile and reviews, but rewrites the answer in its own words, which dilutes the direct link to your listing. The result: a local business well ranked on Google Maps can stay invisible in a ChatGPT answer if its site never explicitly says "we serve Villeurbanne, Vénissieux and Caluire" in text an AI can read. ## How LLMs reconstruct local context An LLM does not know where you are, unless the application tells it. It therefore works with the explicit geographic signals contained in the query and the sources. Three kinds of signals matter. The first is textual mentions of area. A page that lists "Paris, Boulogne-Billancourt, Neuilly-sur-Seine" in full, in a normal sentence, is more likely to be associated with those cities than a page that only puts a postcode in the footer. The second is third-party sources that confirm local presence. A regional press article, a professional directory listing, a dated and localized Google review, a mention on a forum such as Reddit or a local Facebook group. LLMs weigh consistency across several independent sources more than a single site declaring itself the local expert. The third is structured data. LocalBusiness schema is still read by generative engines that crawl raw HTML, provided the page is not entirely client-side JavaScript. An address, opening hours, a service area declared in JSON-LD give the AI a clean signal, with no ambiguity of interpretation. ## What actually makes the difference Many local businesses think one page per city is enough. That is a mistake that already cost dearly in classic SEO, and costs even more in GEO. A "plumber in Villeurbanne" page that copies the "plumber in Vénissieux" content and only changes the city name adds no new information. An LLM comparing several sources detects that duplication and prefers a source that describes a real local foothold: an active local phone number, a named customer testimonial, a recent job told with verifiable details. Content that works answers precise local questions. Not "our plumbing services" but "how much does unblocking a drain cost in Lyon in 2026" or "what is the average plumber response time in the 6th arrondissement". That level of granularity matches what LLM fan-out queries are trying to fill. Local press and specialist directories play a disproportionate role. An article in Le Progrès that mentions your company often weighs more, in an LLM's eyes, than a page on your own site optimized for the keyword. The reason is simple: an independent third-party source is seen as more reliable than a self-declaration. Businesses that invest in local press relations and partnerships with sector directories are, without realizing it, building local GEO capital. Here is how the two disciplines overlap and diverge: | Lever | Local SEO | Local GEO | |---|---|---| | Google Business Profile | Direct ranking signal in Maps | Secondary signal, reused by Gemini especially | | Customer reviews | Influences ranking and click-through | Social proof cited in the generated answer | | Duplicated city pages | Tolerated if slightly differentiated | Penalizing, detected as hollow content | | Local press and directories | Useful for link building | Third-party source the LLM cites for credibility | | LocalBusiness structured data | Optional, helps rich snippets | Key signal if the site is readable without JavaScript | | Real geographic proximity | Direct ranking factor (IP, GPS) | Absent unless the app passes the user's location to the AI | ## Pitfalls to avoid The first pitfall is believing the Google Business Profile is still enough to cover everything. It remains essential for Maps and for Gemini, but it is not read the same way by ChatGPT or Perplexity, which lean more on the open web. The second pitfall is over-optimization by city. Generating fifty automatic pages for fifty neighboring towns produces content that LLMs identify as programmatic spam. Ten pages that tell a precise local reality beat fifty hollow ones. The third pitfall is forgetting that freshness matters. A three-year-old review, a 2021 article, an address that changed without an update: LLMs that prefer recent sources drop these signals in favor of more up-to-date competitors, even modest ones. The last pitfall is treating local GEO as a one-off project. Visibility in generative answers is built over time, by accumulating consistent mentions across the web: a clean site, press, directories, reviews, local social networks. A single audit is not enough. You need monitoring that checks, query after query, how each LLM renders your local presence. *Vurto monitors exactly what ChatGPT, Gemini, Perplexity and Claude answer about your business, city by city, and flags the areas where your local presence stays invisible to AIs.* ### [SEO and GEO: What Changes, What Stays, What to Do Now](https://vurto.ai/en/resources/seo-and-geo-what-changes-what-stays-what-to-do-now) (2026-06-23) There's a convenient temptation in digital marketing: every new discipline gets announced as the death of the previous one. SEO killed display. Social media killed SEO. Now GEO is supposedly killing SEO. That's wrong. And it's a lazy way to think about it. GEO doesn't kill SEO. It changes SEO's role, its scope, and the logic of resource prioritization. That's different. ## What SEO Keeps Doing SEO directly feeds GEO. Well-referenced content — backed by backlinks, well-structured, authoritative in its domain — is what LLMs find first when they search for sources. LLMs don't reinvent the wheel. When Perplexity or ChatGPT go looking for information, they go through Google or Bing. What comes up first influences what they read. What they read influences what they say. A site with strong SEO presence therefore has a solid GEO foundation. It's not sufficient, but it's a real advantage. SEO also remains useful for everything LLMs haven't fully captured yet: local queries with navigational intent, queries where LLMs remain cautious or imprecise, content where users want to explore for themselves rather than receive a synthesis. ## What GEO Does That SEO Doesn't SEO optimizes for a ranking in a list. GEO optimizes for a presence in an answer. The difference sounds semantic. It's structural. To rank on Google, you produce content that matches a search intent, on a keyword, with enough authority and relevance to outrank competitors. The success signal is your position in the results. To be cited in an LLM response, you need to be a source the model judges relevant, credible, and well-structured for a given question. The question can be phrased a hundred different ways. The response the model gives looks nothing like a SERP. And the success signal is the citation, not the position. These two logics coexist. One doesn't replace the other. But they require partially different approaches. Content perfectly optimized for Google can be unreadable for an LLM. Content well-structured in markdown with high informational density can rank modestly on Google and be consistently cited by Perplexity. ## The Priorities That Shift What changes in practice is where attention gets allocated. The race for high-volume keywords becomes less central. Publishing a hundred generic articles on secondary queries to accumulate organic impressions is a strategy that loses value when a growing share of responses to those queries is synthesized by an LLM before reaching the results list. Depth and structure beat volume. A 3,000-word guide that genuinely, precisely, and substantively answers a complex question has more GEO value than a cluster of ten 600-word articles that skim the surface. Diversity of signals matters more. LLMs search publishers, forums, social networks, comparisons. A presence concentrated on your own domain is fragile in GEO. Being present in third-party ecosystems — LinkedIn, Reddit, YouTube, sector press — carries more weight than in SEO. Technical coherence between robots.txt, llms.txt, and LLM crawlers becomes a real issue. It wasn't an SEO dimension. It is a GEO one. ## The Right Framework for Decisions A useful way to prioritize: for each type of content or action, ask whether the goal is to be found by a human browsing, or to be cited by an LLM synthesizing. The two aren't mutually exclusive. Good content can serve both. But when resources are limited, knowing which objective takes precedence for a given piece of content helps calibrate the effort. Category pages, comparison pages, buying guides, expertise pages, data-driven case studies — these are dual-value content types, both SEO and GEO. They deserve full optimization on both dimensions. Short-term news posts, event landing pages, promotional content — their GEO value is more limited. The LLM isn't going to cite a promotional offer in a product recommendation. ## What Vurto Brings to This Transition The difficulty of the SEO-to-GEO transition is that it requires managing two logics in parallel, with partially different signals. Classic SEO tools — Ahrefs, SEMrush, Search Console — cover their perimeter well. They don't see LLM citations, don't measure brand sentiment in generative engines, don't flag content that's unreadable to AI crawlers. Vurto covers that complementary perimeter: citation monitoring in ChatGPT, Gemini, Perplexity, and Claude; content recommendations based on fan-out and grounding queries; AI readability audits; technical file rewriting (llms.txt, robots.txt); and for e-tailers, catalog conversion into structured markdown. It doesn't replace SEO tools. It covers the half of digital visibility they don't see. --- *GEO completes SEO. Vurto covers the part that classic SEO tools don't see.* ### [GEO: The Diagnosis Nobody Is Running](https://vurto.ai/en/resources/geo-the-diagnosis-nobody-is-running) (2026-06-14) SEO has been optimized for thirty years. Rankings, backlinks, Core Web Vitals. All of it worked in a world where traffic flowed through lists of ten blue links. That world is disappearing. Today, a growing share of queries never produces a click. The AI answers directly. ChatGPT, Perplexity, Gemini, Claude — these models synthesize, recommend, conclude. They cite or they don't. And if your brand isn't in the answer, it doesn't exist. That's the territory of GEO. Generative Engine Optimization. ## What GEO Actually Measures SEO measures your position in a list. GEO measures something more fundamental: do AI systems know you, understand you, and recommend you? These aren't the same questions. And they don't have the same answers. A site can rank first on Google and be invisible in ChatGPT's responses. Because LLMs don't read SERPs. They read what they ingested during training, what they fetch in real time through their search tools, and what their architecture lets them process. GEO visibility rests on three layers. The existence layer — do LLMs know you exist? Is your brand name, your products, your core arguments present in the sources these models draw from? The readability layer — are your contents structured in a way that LLMs can process efficiently? A site that's technically opaque to an AI crawler is a site that's transparent to better-calibrated competitors. The perception layer — when AI systems talk about your brand, is what they say accurate, positive, differentiating? Brand perception in LLMs is a signal you can measure, and almost nobody is watching it. ## The GEO Diagnosis, Step by Step The right diagnosis starts with the right questions. Is my brand cited, and in what context? This isn't binary. A brand can be cited positively in Perplexity and absent from ChatGPT. Or present with a negative sentiment in Gemini. Multi-LLM coverage is the first diagnostic layer. Are my contents readable by AI? LLMs process a well-structured markdown document very differently from a proprietary HTML product page or an unindexed PDF. Auditing the AI readability of your key content means finding blind spots before they cost you citations. Do I understand the queries AI systems run about me? LLMs don't answer a user's question directly. They first break it down into sub-queries — fan-out queries — that they send to their search tools. If you don't know which fan-out queries apply to your category, you're optimizing blind. Is my brand perception in AI consistent with my actual positioning? What LLMs say about your brand can diverge significantly from your official messaging. Identifying that gap means knowing where to act first. ## Why the GEO Diagnosis Is Still Rare The short answer: traditional SEO tools don't measure it. Google Search Console gives you impressions and clicks. Ahrefs gives you backlinks. SEMrush gives you positions. None of them tell you what ChatGPT thinks of your brand this morning. The GEO tools market is structuring fast, but maturity isn't uniform. Some platforms do citation monitoring. Others add recommendations. The most advanced ones, like Vurto, cover the full chain: multi-LLM monitoring, brand perception analysis, content recommendations based on fan-out and grounding queries, AI readability checks, and rewriting of strategic content (product pages, llms.txt, e-commerce catalog). The question is no longer whether GEO deserves attention. That's settled. The question is how precisely you run the diagnosis before you start treating. ## What Changes in Practice For SEO teams: the skills transfer partially, but the reference frame shifts. You're no longer chasing a keyword ranking. You're trying to be the source an AI reaches for when a user asks a question related to your category. For content teams: the writing angle changes. Not to please a ranking algorithm, but to be processable by an LLM. Clear structure, informational density, named entities, markdown format where possible. For e-commerce teams: the product catalog becomes a major GEO asset. A poorly structured product page is a product page the AI won't recommend — even if it's perfectly optimized for Google. For marketing leadership: AI visibility becomes a KPI in its own right, with its own benchmarks, trends, and alerts. Not once a quarter at the reporting meeting. Continuously. ## The First Step Run the diagnosis before prescribing. That's true in medicine. It's true in GEO. Start by auditing your site's AI readability. Check whether your llms.txt exists and is properly configured. Ask ChatGPT, Perplexity, and Gemini about your brand and products yourself — what you read there is worth a thousand positioning reports. Then build the monitoring. Not to have one more dashboard, but to know where you stand and where you're heading. --- *Vurto is a GEO tool covering the full diagnostic chain: multi-LLM monitoring, brand perception, AI readability, fan-out queries, and content rewriting for generative engines.* ### Content Strategy ### [ChatGPT, Perplexity, Gemini: four engines, four sourcing logics](https://vurto.ai/en/resources/chatgpt-perplexity-gemini-same-disease-three-different-prescriptions) (2026-09-28) ChatGPT, Perplexity, Gemini and Claude do not draw on the same sources to build their answers. A site visible on ChatGPT can stay invisible on Perplexity. GEO starts by treating one engine at a time. The question often comes back in a form that is too vague: "How do I become visible on AI?" AI does not exist as a single block. There is ChatGPT, Perplexity, Gemini, Claude. Four distinct tools, each with its own way of selecting information. What they have in common: they are chatbots able to answer in natural language by fetching sources on the web, beyond what they learned during training. Their difference, the only one that matters for a brand: they do not draw on the same sources to build an answer. GEO, Generative Engine Optimization (the craft of making a brand visible and well perceived in the answers of these chatbots), starts with a simple observation that is too often skipped. Stop treating these four tools as one uniform set. A site optimized for ChatGPT can be invisible on Perplexity. A brand that shows up strongly on Gemini may never appear on Claude. This is not a bug. It is the direct consequence of different architecture choices, made by different companies, with different interests. ## Four engines, four sourcing logics Step back and look at how each engine picks its sources. Independent studies run on tens of thousands of citations in 2026 draw very sharp, almost caricatural profiles. Perplexity behaves like a reporter who favors the field. It relies heavily on YouTube (around 30% of its citations in some analyses) and on Reddit, the forum platform where people discuss everything, often more candidly than a press release. Reddit can account for nearly half of its top ten sources. Perplexity prefers lived experience told by a stranger over the claim of an institutional website. ChatGPT, for its part, leans toward established authority. Wikipedia holds a disproportionate place, and well-known media (Forbes, Reuters) complete the picture. The profile looks like a serious student who cites academic sources rather than comments from the neighborhood forum. But watch what comes next, because this profile moves fast. Gemini, owned by Google, keeps the reflexes of a corporate librarian. It keeps citing the group's own properties widely (YouTube first) while keeping Reddit and Wikipedia in a good spot. A more balanced split, consistent with an engine that has access to the most complete search index on the web. Claude, developed by Anthropic, completes the picture with yet another profile, closer to ChatGPT in its taste for documented sources, but giving a larger share to technical documentation and specialized sites when the question calls for it. Four engines, four sourcing logics, and no single strategy that fits all. A table beats a long speech here. | Engine | Dominant sources | Profile | |---|---|---| | Perplexity | YouTube, Reddit | Field, lived experience, community | | ChatGPT | Wikipedia, recognized media | Authority, formal knowledge, volatile | | Gemini | YouTube, Reddit, Google properties | Balanced, anchored in the Google ecosystem | | Claude | Documentation, specialized press, Wikipedia | Technical, precise, expertise-oriented | Remember one thing from this table. Nobody builds an effective content strategy by targeting "AI" in general. You treat one engine at a time. An HR software brand that neglects Reddit cuts itself off from Perplexity. A services company that never keeps an up-to-date Wikipedia page loses part of the credit it would deserve on ChatGPT. These are not technical details reserved for specialists. These are choices that decide which of your competitors gets recommended instead of you. ## When an engine's sources shift in a few weeks Here is the most underestimated point of the topic, and probably the most important. These profiles are not set in stone. A study covering 230,000 queries tracked over thirteen weeks showed that Reddit's share in ChatGPT answers went from about 60% in August to only 10% six weeks later. Wikipedia, in the same movement, dropped from 55% to under 20%. A collapse of that scale, over such a short time, on one and the same platform. Meanwhile, Perplexity stayed remarkably stable, with minor variations on its usual sources. Gemini, too, changed little. This contrast says something essential about these tools. ChatGPT probably adjusts its weightings continuously to avoid over-citing a handful of domains and looking like a biased engine. Perplexity and Gemini, with different sourcing logics, absorb these adjustments better. The message for you: a position won on one engine in September can vanish in October, without you having made any mistake. The ground moves under your feet. One more reason never to talk about "AI positioning" as a permanent gain, but as a state to monitor continuously. ## What this concretely changes for your content This heterogeneity has a direct consequence on how you should produce and distribute your content. Stop betting on a single channel. If your content strategy is limited to blog articles optimized the classic search way, you may cover part of ChatGPT, but you stay invisible on Perplexity. An active presence in Reddit discussions relevant to your sector, YouTube videos that explain your product, an up-to-date Wikipedia page if your notoriety justifies it: these are different building blocks, for different engines, and none replaces the others. Also beware of wins that rest on a single source. A brand that owes all its AI visibility to a handful of Reddit mentions lives on unstable ground, as we just saw with ChatGPT. Robustness comes from the diversity of entry points, forums, video, specialized press, well-structured product pages, kept up to date at the same time rather than one isolated bet. And above all, do not confuse being cited with being well perceived. Appearing in an answer guarantees neither a favorable tone nor an accurate description of your offer. That is a topic in its own right, but it starts with the same prerequisite: knowing precisely where and how you appear, engine by engine, rather than guessing. Take a concrete example. A company that sells software to businesses (what is called B2B, as opposed to selling directly to consumers) usually invests heavily in white papers and case studies published on its own site. This content feeds ChatGPT well, especially if it is picked up by the specialized press. But it barely feeds Perplexity, which will rather look for a customer's opinion posted on a forum or a commented LinkedIn post. Two audiences, two formats, two engines. The content budget must reflect that reality, not ignore it. ## Measure before you act The temptation, faced with this complexity, is to give up and tell yourself that "AI moves too much to be worth our time". That is the reverse of the mistake we pointed out at the start. Movement does not exempt you from monitoring, it demands it. A serious company does not decide its sales strategy on an intuition from six months ago. It looks at this month's numbers. The logic is the same here. What matters is knowing, for your brand and your competitors, which engines cite you, which sources they rely on to do so, and how that split evolves over time. Without that regular snapshot, you optimize blind, hoping that what worked on ChatGPT in the spring still works on Perplexity in the fall. That is a bet, not a strategy. *Vurto tracks these variations engine by engine, ChatGPT, Gemini, Perplexity and Claude, to show precisely which sources carry your visibility and which ones are collapsing before you find out by chance.* ### [Perplexity, ChatGPT, Gemini: Three Engines, Three Source Lists, One Mistake to Avoid](https://vurto.ai/en/resources/perplexity-chatgpt-gemini-three-engines-one-mistake-to-avoid) (2026-09-14) A number to start with. Out of 680 million citations analyzed between August 2024 and April 2026 by the 5W index, only 14% of cited sites dominate across all AI chatbots (these conversational assistants like ChatGPT, Gemini, or Perplexity, able to answer in natural language from sources they select themselves). Between ChatGPT and Perplexity alone, overlap drops to 11%. In another study covering 11,500 queries compared against Google results, GPT-4o shows 0% median overlap with Google's top 10. Perplexity, 14.3%. Gemini, 8.5%. These numbers should settle a debate that has dragged on for two years in marketing teams: no, "being visible in AI" is not a single objective. It's three different objectives, with three different selection logics, and a good part of the groundwork consists of understanding which one applies to which client. ## Three philosophies, not three variants of the same engine The starting mistake is to treat ChatGPT, Perplexity, and Gemini as competitors doing the same thing with different branding. They are three architectures that answer different questions about what makes a good source. ChatGPT relies heavily on consensus-based, already-aggregated content. Wikipedia accounts for between 26% and 48% of its top 10 citations depending on category. Reddit comes right behind, with a weight exceeding 40% across all categories and engines combined, and staying particularly high with ChatGPT. The logic is that of a model favoring what has already been collectively validated: content synthesized, debated, corrected by thousands of contributors. Perplexity works the opposite way. Its engine queries the live web on every request rather than drawing from a fixed memory, which explains why it favors Reuters, AP, Bloomberg, the Wall Street Journal: news sources with strict attribution and verifiable freshness. A Columbia Journalism Review study measured a citation error rate of 37% for Perplexity Sonar Pro, versus 67% for ChatGPT Search. The gap comes down to method: searching live and citing your source produces fewer fabrications than rephrasing from training memory. Gemini, finally, remains Google's child. It draws from the search engine's index, from YouTube (which Google owns), and shares with AI Overviews a clear fondness for Reddit. It's the engine closest to classic SEO culture, the one where web ranking, domain authority, and presence across the Google ecosystem still count directly. Claude, often looked at from a distance in these comparisons for lack of search volume, stands out for a taste for long-form analysis and technical documentation. It cites proportionally more expert blogs and less Reddit than the others. ## The table that changes the roadmap | Engine | Preferred sources | Dominant logic | |---|---|---| | ChatGPT | Wikipedia, Reddit, consensus content | Trained memory + collective synthesis | | Perplexity | Reuters, AP, Bloomberg, press with attribution | Real-time search, freshness | | Gemini | Google index, YouTube, Reddit | Google ecosystem, classic SEO | | Claude | Expert blogs, technical documentation | Long-form analysis, depth | This table isn't an exercise in curiosity. It dictates where a brand needs to exist depending on which engine matters most to its audience. A B2B software vendor (selling to other businesses rather than individuals) whose prospects mainly use Perplexity to compare solutions is better off working on press relations and citations in serious trade media, far more than optimizing a Reddit thread. A consumer brand that lives and dies on ChatGPT has the opposite interest: community conversation outweighs the press release. ## What the low overlap actually means The most underestimated point isn't the list of preferred sources. It's the fact that they barely overlap at all. Content that breaks through on ChatGPT can be invisible on Perplexity, and vice versa. This means a content strategy designed for "generative AI" in general, without distinguishing between engines, is in reality optimizing for a single engine without realizing it, usually the one the brand heard about first, often ChatGPT because it concentrates most of the usage volume. In France, this concentration reaches extreme levels: ChatGPT captures 84.47% of clicks generated by AI assistants according to SE Ranking, Perplexity 12.82%, Gemini 2.08%. An AI visibility budget that ignores this split and treats the three engines equally wastes precious time. But the reverse is also true over an eighteen-month horizon: Gemini is natively built into Android, into Google Workspace, into the Google search bar itself. Its share of usage is growing faster than its share of citations today would suggest. Ignoring Gemini because it accounts for 2% of current clicks is like ignoring mobile in 2010 because it still represented only a fraction of web traffic. The practical consequence fits in one sentence: stop measuring "AI visibility" as a single score. Content can be cited ten times a week on ChatGPT and never on Perplexity without that being a failure, if the brand's target audience lives on ChatGPT. But no one can know that without tracking the engines separately, each with its own citation logic, rather than an average that hides everything. ## Adapting content to each engine's logic, not to a universal checklist Concretely, this changes three things in how content gets produced. To exist on ChatGPT, you need to feed the collective conversation rather than bypass it. A structured presence on Reddit in the right communities, an up-to-date Wikipedia entry if the brand or its leader has a place there, content that synthesizes a topic rather than selling it. It's ecosystem work, not landing-page work. To exist on Perplexity, freshness and attribution matter more than volume. A well-distributed press release, relationships with journalists who cite the brand with a clear link, dated content republished regularly carry more weight than a static white paper published once and never touched again. To exist on Gemini, classic web ranking fundamentals remain the best way in: domain authority, clean technical structure, YouTube presence. It's the engine where the last ten years of SEO work keep paying off directly, which paradoxically makes it the easiest to work on for a team coming from classic search. None of these three approaches replaces the other two. A brand with the time and budget to do all three builds robust AI visibility, insensitive to a single engine's algorithm changes. A brand that has to choose should first look at where its audience actually asks its questions, not where it itself is most comfortable producing content. ## Tracking three engines beats guessing on one The low overlap between ChatGPT, Perplexity, and Gemini isn't a technical problem to solve. It's a structural constraint to build into how you measure and prioritize. A brand that only tracks its citations on one engine is flying blind on the other two, and makes content decisions based on a third of reality. *This is exactly why Vurto tracks citations separately across ChatGPT, Gemini, Perplexity, and Claude rather than producing one aggregated score that erases these differences: a brand needs to know which engine it exists on, which one it's absent from, and why the two don't look alike.* ### [Google's AI Summary That Killed the Click](https://vurto.ai/en/resources/googles-ai-summary-that-killed-the-click) (2026-09-10) Type a question into Google today, and odds are a box appears right at the top of the page, ahead of the usual list of links. That box is called an AI Overview, literally an "AI-generated overview." A summary written by Google's AI model that answers the question directly, with a few small links off to the side. The person searching for information already has it in front of them. They no longer need to click on anything. This box now appears on 15 to 25% of all Google searches, and up to half of searches when the question is informational rather than transactional in nature. Its direct consequence shows up in a number that should alarm any site that lives off Google traffic: when this summary appears, the click-through rate to the sites ranked below it drops by nearly 60%. A good chunk of the free traffic that thousands of businesses have built their visibility on for fifteen years is evaporating before their eyes, replaced by text that Google writes itself from their content, without sending them a single visitor in return. ## A site can be the source and receive no visitors The most disorienting part of this mechanism is that a site can supply the information shown in Google's summary without a single visitor ever clicking on it. On searches where this type of summary appears, 83% of people close the page without clicking on any link, neither the one in the summary nor those in the classic list. On Google's more advanced search mode, where AI takes up an even bigger share of the page, that figure climbs to 93%. More broadly, the share of Google searches that end with no click at all rose from 50% in 2019 to nearly 65% in early 2026, and reaches 68% in the United States over that same period. This shift changes the very definition of what it means to "be visible" on Google. For twenty years, being well-ranked meant showing up at the top of a list of links and getting a click. Today, you can be the source cited in a summary, read by thousands of people, and never know it, because no visit shows up in the site's statistics. Visibility still exists, but it no longer automatically translates into measurable traffic. Take a site that has published a detailed guide on boiler maintenance for years. It held the top spot on Google and got several hundred visits a day on that one page alone. Google's generated summary now captures the gist of that guide in a few lines, with a discreet link back to the source. The site's content answered thousands of people's questions this month. The site itself saw only a few dozen extra visits. The editorial work hasn't lost its value to the reader, it has simply changed who visibly benefits from it. ## What determines who gets cited in the summary Google's generated summary works on a logic close to that of chatbots like ChatGPT or Perplexity. The model has to choose, among the pages that best answer the question asked, which ones to cite and summarize. A poorly structured page, filled with sales pitches before getting to the point, or one that requires a lot of interaction to display its content, has less chance of being chosen than a page that answers directly and clearly from the very first lines. Three elements come up in most analyses of pages cited by these summaries. First, an answer formulated in a self-contained, direct way, one that makes sense even taken out of its context, since that's exactly what the model is going to do with it. Second, a clear structure, with headings that announce the content that follows, rather than one long continuous block of text where the information gets buried. Third, a perceived authority on the topic covered, built through the consistency and regularity of the content published on that theme, not through a single article, however well written. | What helps you get cited | What limits your chances of being cited | |---|---| | Clear answer from the first lines | Sales pitch before the answer | | Headings that structure the information | Continuous text with no visual landmarks | | Content accessible without interaction | Content loaded only through complex scripts | | Consistency over time around a topic | A standalone article, with no related content on the same theme | ## The clicks that remain are worth more There's a paradox worth noting before giving in to discouragement. While the presence of a Google-generated summary causes clicks to drop, the people who click anyway convert 23% better than on a classic search. These visitors have already read a summary of the answer. If they click through to a site on top of that, it's because they're looking for more precise information or want to take action, not because they're just starting their search. The mechanism resembles what's been observed with traffic sent by AI chatbots like ChatGPT: less volume, but a visitor who is much further along in their decision. That doesn't offset the loss in volume for most sites, but it does change how you should evaluate your content strategy. A site that loses 60% of its clicks on a query but gains citations in the summaries isn't necessarily losing influence. It's losing old-style measurable traffic, while gaining exposure that builds its reputation among people who will never open its site but will remember its name, because Google presented it to them as a trustworthy source. ## Adapting rather than enduring Faced with these numbers, the temptation is to conclude that producing content no longer serves any purpose since the click is disappearing. That's the wrong conclusion to draw. Content matters more, not less, because it now has to convince a model to choose it before it can convince a human to click. The difference is that the format has to change: less sales pitch up front, more direct and structured answers, a consistent presence on the topics where you want to be recognized as a reference. The real risk isn't losing traffic on Google. It's continuing to write for a human reader who skims a page before clicking, when the first reader of much of that content is now a model deciding, in a few seconds, whether it's worth citing. *Vurto checks whether your content is genuinely structured to be understood and cited by these summaries, on Google as well as on AI chatbots, before the traffic drop becomes hard to explain internally.* ### [GEO for B2B Brands: Why Your Prospects No Longer Read Reviews Themselves](https://vurto.ai/en/resources/geo-for-b2b-brands-why-your-prospects-stopped-reading-reviews-themselves) (2026-09-07) GEO — the art of being visible in the answers given by chatbots like ChatGPT, Claude, or Gemini rather than in Google's blue links — doesn't play out the same way depending on who's buying. A brand selling to individual consumers and a brand selling to other businesses (what's known as B2B, for "business to business") aren't playing on the same field. The latter is losing control of its sales funnel without even realizing it. In April 2026, G2, one of the largest online review platforms for business software, published a number that should alarm any B2B marketing team. 51% of software buyers now start their research in an AI chatbot, up from 29% a year earlier. ChatGPT alone captures 63% of that activity. Over the same period, human traffic to G2's site dropped 84.5%, its competitor Capterra's by 89%, and TrustRadius's by 92%. These three names are the reference sites where businesses publish and read reviews on business software, the B2B equivalent of a specialized Trustpilot. A study by 6sense of nearly 4,000 B2B buyers confirms the trend on a larger scale: 94% of them use an AI chatbot at some point in their buying journey. These numbers don't tell the story of review sites dying. They tell the story of a shift. Buyers aren't giving up on reading customer reviews, they're giving up on going to read them themselves. A chatbot does it for them, synthesizes, compares, and recommends three tools before a salesperson has even known a prospect existed. For a B2B brand, the question is no longer whether it shows up on the first page of Google. It's whether it's one of the three names ChatGPT cites when a buyer asks "what's the best CRM software for a team of 20 salespeople." ## A buying cycle that starts outside your control B2B selling has always had a long buying cycle, with several decision-makers and a silent research phase before any sales contact. GEO doesn't create this dynamic, it moves it further upstream. What used to happen on Google, on comparison sites, and on specialized forums now happens in a conversation with a chatbot, invisible to you, leaving no trace in your CRM, nothing to observe in your site statistics. A direct and underestimated consequence: 69% of buyers surveyed by G2 say they chose a different vendor than the one they initially had in mind, based solely on a chatbot's recommendations. A third bought a tool they had never heard of before the conversation. The shortlist a prospect builds before ever talking to a salesperson is no longer the product of their own experience or word of mouth. It's the product of a synthesis produced by a chatbot, from sources it deems reliable. If your brand isn't among those sources, you're not just losing a click. You're losing the sale before it ever existed. Take an operations director looking for inventory management software for an industrial SME. Two years ago, they would have typed their query into Google, clicked on three or four results, compared pricing grids across ten open tabs. Today, they ask ChatGPT the question in one sentence, get three names with their strengths and limitations, and only contact the ones that come out of that synthesis. The salesperson for the software left out of that answer will never know an opportunity existed. No form filled out, no trace in their CRM, just a silent absence in a conversation no measurement tool can capture. ## G2, Capterra, TrustRadius: the trust base behind AI chatbots Here's the paradox many marketing teams haven't yet come to terms with. Human traffic to review sites is collapsing, but their weight in AI chatbot answers is growing. A Quoleady study of near-purchase searches, of the "[software] alternatives" type, found that 100% of recommended tools had reviews on Capterra and 99% on G2. These platforms are no longer read by humans, they're being digested by models as aggregated social proof. A chatbot treats a review published on G2 as a more reliable signal than a product page, because it carries the mark of an independent judgment. A report published by G2 in early 2026 goes further: the platform's recent acquisition could push its share of citations in pre-purchase answers up by 76%, simply because the volume of reviews it aggregates becomes strategic raw material for the models. For a B2B brand, ignoring its presence on these platforms while betting everything on its own site is a miscalculation. The corporate website remains useful for convincing a prospect once they've arrived. It carries almost no weight in the phase where the chatbot builds its recommendation. ## What sets selling to businesses apart from selling to consumers Consumer-facing GEO, the kind aimed at individuals (B2C, "business to consumer"), tries to get a product into a one-off recommendation: what's the best sunscreen, which vacuum to buy. GEO aimed at businesses plays a different game. It tries to build perceived authority across an entire category, because the final decision involves several people, several research sessions, and a comparison that stretches over weeks. | Dimension | Selling to consumers | Selling to businesses | |---|---|---| | Goal | Show up in a one-off recommendation | Be perceived as a reference across the whole category | | Sources that matter | Reddit, product reviews, consumer comparisons | Professional review sites, technical documentation, expert publications on LinkedIn | | Decision | Fast, often a single person | Long, several decision-makers to convince | | Content that carries weight | Short reviews, social content | Detailed comparisons, case studies, technical proof | Three levers matter in particular when selling to businesses. First, the depth of technical documentation: a chatbot that has to answer "does this tool connect with Salesforce" will look for a precise integration page, not a marketing pitch. Second, honest comparative content: pages that fairly compare a product to a competitor, without hiding its limitations, get cited more often than pages that only praise their own merits, because they answer directly to the implicit question of a buyer in the selection phase. Third, the credibility of the company's experts, on LinkedIn in particular, which strengthens the trust models place in a source before citing it. A B2B buying committee averages six to ten people, each with their own criteria. The technical lead is looking for an answer on security and integrations with other software. The finance lead wants to understand total cost over three years. The future end user wants to know if the tool is pleasant to use day to day. A chatbot queried by each of these profiles will look for different sources to answer different questions. A brand that only covers the product angle, without solid security documentation or clear financial content, loses citations on half the questions its own buying committee is asking. ## Invisibility now costs more than it used to A third of B2B buyers today buy a tool they didn't know about before their conversation with a chatbot. That number needs to be read both ways. It's a threat to established brands that thought their position was secured by their historical name recognition. It's a real opportunity for smaller or newer players, who can earn a spot in an AI recommendation without a market leader's budget, provided they've built the right signals: structured presence on review platforms, clear and accessible documentation, honest comparative content, identifiable expert voices. The real risk isn't doing GEO badly. It's continuing to measure performance only with yesterday's tools. A team that tracks its traffic to G2 and reads its decline as a sign of waning market interest is missing the opposite signal: the market is still reading those reviews, just through a chatbot intermediary. The question to ask every quarter is no longer "how many visitors on our product page," it's "when a buyer in our category asks ChatGPT or Perplexity, are we in the answer." *Vurto tracks this question precisely for B2B brands: which queries generate a citation, which sources carry weight in the answer, and where the gaps are against the competition.* ### [GEO and Video Content: YouTube, the New Preferred Source of Generative Engines](https://vurto.ai/en/resources/geo-and-video-content-youtube-new-preferred-source-of-generative-engines) (2026-08-31) Between August and December 2025, Reddit's share of social citations in AI answers fell from 44.2% to 20.3%. Over the same period, YouTube's jumped from 18.9% to 39.2%. An Adweek survey published in January 2026, combining data from four research firms across 6.1 million citations, confirms the shift. YouTube now ranks ahead of Reddit as the social platform most cited by ChatGPT, Perplexity and Google AI Overviews. The forum that had dominated citations for two years has just lost first place. Not to an unexpected competitor, but to a platform most marketing teams still treat as a distribution channel, never as a knowledge source. That is the first mistake to correct. ## Why a video becomes a citable source An LLM does not watch a video. It reads its transcript. That is the starting point to understand before anything else. When a generative model cites YouTube, it is actually citing a text: the automatic or manual captions of the video, enriched by the description and chapter metadata. The video itself remains invisible to the model. Its transcript becomes a document like any other in the corpus the LLM queries. This mechanic explains why YouTube weighs so much. The datasets used to train the large models, HowTo100M and Kinetics first among them, rest largely on YouTube transcripts and metadata. The models grew up with this platform as a reference. It is familiar to them in a way no other video platform has managed to match — not TikTok, not Vimeo, not videos hosted on a brand site. There is a second, more structural reason. A well-produced video condenses an oral, precise, often demonstrative answer to a question someone is actually asking. A tutorial that shows how to configure a tool, an expert who walks through a method step by step, a designer who comments on their choices in real time: this format produces a source text different from a blog post. More direct, closer to spoken language, often more honest about limits and edge cases. LLMs value this kind of raw material when building an answer that sounds right. ## What makes a video actually usable by an LLM Not all videos are equal in front of this mechanism. A video with no clean transcript, no chapters, and a three-word description remains almost invisible to a generative engine, whatever its view count. Three technical elements determine a video's real citability. | Element | Role for AI citation | |---|---| | Transcript | The text the LLM actually reads. Without it, no citation is possible, regardless of content quality | | Chapters | Let the engine point to a precise timestamp rather than the whole video | | Structured description | Serves as a usable summary; must answer the question in the first two sentences | The transcript comes first because without it, nothing else matters. An auto-generated transcript with poor punctuation and mangled proper names produces a low-quality source text that the model either avoids or distorts when citing it. A human pass, even a quick one, fundamentally changes the reliability of that raw material. Chapters come next. A model able to point to the exact minute where you answer a question gains citation precision, and that precision becomes a trust criterion. A twelve-to-fifteen-minute video split into five to seven chapters of two to three minutes each offers a granularity that matches the level of detail expected in a generative answer. The description, finally, should stop being an SEO formality stuffed with keywords. It works better treated as a short two-hundred-word article that answers the question directly in the opening lines, before detailing the content and timestamps. It is often this text, more than the full transcript, that engines synthesize first to build a short citation. ## What this changes for a B2B content strategy The classic mistake is to treat YouTube as one social channel among others, on the same level as LinkedIn or Instagram, with a views-and-engagement objective. That reading misses the point. YouTube now functions as a knowledge base that generative engines query on the same footing as a website or technical documentation. For a B2B company, that opens largely underused ground. A recorded webinar, a narrated product demo, a Q&A session with a client: these are formats that already exist in most marketing teams, often published with no chapters and no careful description, sometimes even without subtitles. Fixing those three points on content already produced costs little and turns hours of dormant video into citable material. The most profitable shift is not to produce more videos. It is to take what already exists and make it readable for an LLM. A conference filmed last year, an internal tutorial never republished, a client exchange captured at an event: every piece of that catalog deserves a quick audit before anyone invests in new production. The competitive window is still wide right now, with most SEO teams still treating YouTube as a branding channel rather than a citation source. ## The traps that cancel the effort Two mistakes come back most often. The first is leaving the automatic transcript as-is on technical content or industry jargon. YouTube transcribes acronyms, product names and specialized terms poorly. An LLM that reads a transcript riddled with errors on your own vocabulary cannot cite you correctly — and worse, it can cite a distorted version of what you said. The second mistake is treating chapters as a YouTube constraint rather than a tool for structuring the argument. Vague chapters — "introduction", "part 1", "conclusion" — add nothing. Chapters phrased as precise questions or statements — "how to calculate the ROI of a GEO audit", "the three errors that break a transcript" — give the model a clear anchor to extract an answer and tie it to an exact timestamp. There is a quieter trap as well: thinking that video replaces written content. It does not replace it; it complements it. A generative engine builds its answers from a set of converging sources. A well-structured video, an article that covers the same topic in depth, a clear product page: together, these formats reinforce the same authority signal on a given subject. Isolated, a perfectly optimized video remains one source among others, not a shortcut to visibility. ## What to remember before shooting the next video YouTube is no longer one distribution channel among others for a brand thinking about its AI visibility. In a few months it has become the leading social source cited by generative engines, ahead of Reddit which had dominated until now. The change does not require a higher production budget. It requires new discipline on three simple technical elements: a clean transcript, precise chapters, and a description that answers the question before describing it. Teams that adjust these three points on their existing catalog, before even producing new content, take a lead that compounds over time. Those that keep treating YouTube as a mere view counter leave a growing share of their AI visibility to competitors — or worse, to third parties who talk about their product without ever having been invited. *Vurto's multi-LLM monitoring tracks your brand citations on ChatGPT, Gemini, Perplexity and Claude, YouTube included when a transcript or a comment mentions your name. So you can tell whether your videos are actually working for your AI visibility, or sitting unread.* ### [Reddit and GEO: Joining Conversations Beats Posting Into the Void](https://vurto.ai/en/resources/reddit-and-geo-joining-conversations-beats-posting-into-the-void) (2026-08-03) A number to start with. Reddit now accounts for about 40% of citations across ChatGPT, Claude, Perplexity and Google AI Overviews combined, ahead of Wikipedia and YouTube put together. On Perplexity alone, some March 2026 studies put Reddit's share as high as 46.5% of citations. Growth over the last four months exceeds 70% in the categories tracked. No other domain comes close to this level of presence in generated answers. For a brand that thinks of AI visibility in terms of a blog, product pages or LinkedIn posts, that number should set off an alarm. The most-cited channel by generative engines is not an editorial channel you control. It is a forum where strangers argue, contradict each other, and where your brand has historically had no organized voice. Understanding why changes the strategy from the ground up. ## Why LLMs prefer Reddit to brand content Generative models look for answers that sound like what a competent human would tell a friend. Reddit produces exactly that kind of text, in millions of threads, on every topic imaginable. A thread on "which CRM to pick for a ten-person sales team" contains conflicting opinions, dated first-hand accounts, real-time corrections when someone gets it wrong. It is a corpus of argued disagreement, exactly what LLMs synthesize best in order to give a nuanced answer rather than a single, suspicious take. Perplexity is especially calibrated for this. The engine looks for real people answering real questions, not pages optimized to rank. A Reddit comment with twelve upvotes and a well-built counter-argument carries more weight than a landing page claiming to be "the best solution on the market" without ever saying why. Gemini, by contrast, cites Reddit in barely 0.1% of its answers. The gap between engines matters as much as the overall trend: your Reddit strategy will not pay off the same way depending on which LLM your customers actually use. There is a contractual reason behind these numbers as well. Google pays sixty million dollars a year to Reddit for API access to its data and to train its models on it. OpenAI signed a comparable deal, around seventy million a year. Those deals are being renegotiated during 2026, with Reddit trying to index the price to the real value its data brings to AI answers. Whatever the outcome of those negotiations: they confirm that LLMs treat Reddit as a primary source, not as an accessory forum they crawl at random. ## Publishing a post is not enough — participating changes everything The most common mistake is to treat Reddit like LinkedIn: create a brand account, post content regularly, hope it takes. It does not work. Reddit culture rejects content that smells like promotion from ten meters away, and communities have increasingly fine-grained detection mechanisms for manufactured engagement. An account created yesterday that posts a link to its own site gets banned within hours in most serious subreddits. What works is the opposite: reply in existing threads where your expertise has a legitimate place. Someone asks "how do we know if our site is actually readable by generative AIs" in a marketing or SEO subreddit. A precise answer, with a concrete example and a verifiable number, builds an expertise signal that neither a blog post nor an ad can reproduce. That comment becomes a potential source for an LLM that later synthesizes an answer to a similar question, with or without a direct citation of your brand. This logic inverts the usual content marketing hierarchy. On a blog, you control the topic, the calendar, the message. On Reddit, you answer real demand, formulated by someone with a precise problem, at the moment they formulate it. It is slower to build, less predictable, but infinitely closer to what LLMs are trying to extract: a useful answer to a question asked by a human, not content produced to exist. ## The self-promotion trap and how to avoid it Every subreddit has its own rules, usually displayed in the sidebar or a dedicated wiki. Some ban commercial links outright. Others tolerate a mention if it is clearly identified as such and brings real value beyond the link. Ignoring these rules does not just result in a post being removed. It results in the account being banned, and sometimes in an entire thread turning against the brand, which stays indexed and visible long after. The survival rule fits in one sentence: observe before you act. Spend time in a community before speaking there, understand its tone, its codes, what is done and not done. A B2B brand that shows up in a technical subreddit with a commercial tone gets spotted immediately, and the social sanction on Reddit is fast and public. Transparency remains the best lever available. An employee posting under their real account, stating who they are and why they are answering, lands much better than an anonymous brand account pretending to be a neutral user. Reddit tolerates assumed expertise. It does not tolerate disguise. | Approach | Likely result | Why | |---|---|---| | Brand account, regular promotional posts | Fast ban | Self-promotion detection, cultural rejection | | Identified employee, occasional useful replies | Growing credibility | Assumed expertise, real value | | Fake user account, discreet steered reviews | Risk of backlash | Community detection, boomerang effect | | Transparent sponsoring of an AMA or a post | Decent, controlled visibility | Accepted format if clearly disclosed | ## What this changes concretely for your GEO strategy Reddit is not a channel to add at the end of the list after the blog and LinkedIn. Given its weight in citations, it is a channel to treat with the same rigor as any other strategic source, with a specific angle: individual human presence outweighs collective brand presence. Concretely, that means identifying the subreddits where your customers actually ask their questions, not the ones that look relevant on the surface. It means designating one or two internal people, with real expertise, who can reply regularly without it looking like a forced marketing task. And it means tracking what happens: which Reddit threads surface in ChatGPT or Perplexity answers on your strategic queries, which comments become cited sources, and which stay invisible despite the effort. Without that tracking, Reddit presence remains a blind bet. With it, it becomes a measurable lever, on the same footing as SEO or advertising. The difference with more classic channels is that on Reddit, authenticity is not a stylistic option. It is the only strategy that works. *At Vurto, multi-LLM monitoring lets you see exactly when a Reddit thread or comment becomes a source cited by ChatGPT, Gemini, Perplexity or Claude on your strategic queries, so you can adjust where and how your expertise should show up.* ### [Grounding Queries: The Hidden Layer That Decides Your AI Visibility](https://vurto.ai/en/resources/grounding-queries-the-hidden-layer-that-decides-your-ai-visibility) (2026-06-19) Fan-out queries are what the LLM sends outward to search for information. Grounding queries are different. They're the internal signal by which the model identifies what it doesn't yet know — before it even starts searching. Understanding the difference between the two is understanding how an LLM decides what to go looking for — and therefore which content it ends up pulling. ## The Internal Mechanics of an LLM Answering a Question When a user asks a question to an LLM with web access, the process works roughly like this. The model first evaluates whether its internal knowledge base — what it learned during training — is sufficient to answer. If the response can be built from stable, well-documented facts, the model answers directly. If the question touches on recent information, changing data, or updated comparisons, the model identifies those gaps. Those are the grounding queries. The model then formulates search queries to fill those gaps — those are the fan-out queries. It retrieves the results, synthesizes them, and builds the final response. Grounding queries are the upstream layer. They determine what the model knows it doesn't know. And therefore what it goes looking for. ## Why Grounding Queries Are a GEO Lever If an LLM considers your brand, product, or category to be part of its stable knowledge, it won't go looking for additional information. It answers from its internal base — which may be outdated or incomplete. If the model identifies grounding queries about your brand (recent pricing, new features, current reviews), it searches. And then the quality and structure of what it finds determines what it says. This lever cuts both ways. Favorable case: the model looks for recent information about your brand and finds your well-structured content, up-to-date comparisons, readable product pages. The response it builds is precise and positive. Unfavorable case: the model searches and finds an old comparison that ranks you poorly, an article mentioning a bug you fixed a year ago, or nothing structured about you at all. The response is imprecise, dated, or absent. ## The Types of Grounding Queries That Apply to Your Brand Brand-related grounding queries generally fall into these categories. Recent news. "Have there been any recent developments about [brand]?" The model knows its training data has a cutoff. For anything that may have changed, it searches. Pricing and offers. Prices change. Plans evolve. Models know this and systematically look for pricing information in real time from comparison sites and pricing pages. User reviews. Reputation evolves. Recent reviews are a frequent grounding query for purchase-intent questions. Comparisons. "How does [brand] compare to [competitor] today?" Recently published third-party comparisons are a commonly mobilized source for these queries. New features. If your brand launched a feature after the model's knowledge cutoff, that's a gap the model may try to fill if it knows the brand is evolving. ## How to Work on Grounding Queries The strategy plays out on two fronts. First: make sure the model finds recent, reliable information. That means maintaining a consistent flow of content on the subjects that grounding queries cover — pricing clearly displayed on a dedicated page, product updates documented with a clear date, recent reviews accessible and crawlable, up-to-date comparisons. This isn't different from good SEO hygiene. But the calibration shifts: the goal isn't to rank, it's to be the source an LLM finds when it's trying to validate a piece of information. Second: optimize the structure of that content for grounding queries. A pricing page with a clean, dated, structured table is more usable by an LLM than a pricing page with interactive cards loaded in JavaScript. A "Vurto Updates — June 2026" article with clear, dated bullet points directly answers grounding queries about recent changes. ## The Difference from Classic Content Marketing Classic content marketing produces content to attract human readers on high-volume queries. The grounding query logic is different. The goal isn't to drive traffic with this content. It's to make it available at the moment an LLM is trying to validate a piece of information about your brand. The traffic these pages generate doesn't matter if nobody reads them directly. What matters is that they exist, are well-structured, crawlable, recent, and aligned with what the model is looking for. It's a logic of information stock rather than traffic flow. ## The Grounding Query Audit A grounding query audit on a brand answers these questions. On which topics do LLMs search for additional information when they process a query related to your brand? Which sources do they find? Are those sources recent, favorable, well-structured? Are there gaps — subjects where the model searches and finds no clear signal about your brand? Vurto integrates grounding query analysis into its content recommendation process. The tool identifies the blind spots — topics where LLMs search and find no clear signal about your brand — and recommends which content to create or update first. ## What You Learn by Looking at Grounding Queries The first lesson, consistent across brands: pricing. Almost always a grounding query. Almost never a pricing page that's readable by LLMs. The second: third-party comparisons carry more weight than your own pages. When the model tries to validate an advantage or disadvantage of your brand, it prefers a source perceived as independent. The third: new features disappear into silence if they're not documented in a crawlable format with a clear date. Addressing these three points covers most of the grounding query work for the majority of brands. --- *Vurto identifies the grounding queries associated with your brand and recommends content to create or update to fill detected gaps.* ### [Query Fan-Outs: The Blind Spot of 99% of Marketing Teams](https://vurto.ai/en/resources/query-fan-outs-the-blind-spot-of-99-of-marketing-teams) (2026-06-15) Here's what happens when you ask ChatGPT "what's the best GEO tool in 2026." The model doesn't answer directly. It first breaks your question down into a series of sub-queries it sends to its search tools — Google, Bing, or its own indexes. These sub-queries are what we call query fan-outs. The result: it's not your well-ranked pages on "best GEO tool" that determine your presence in the response. It's the pages and sources the model finds across each of its decomposed sub-queries. And nobody is watching that. ## Why Fan-Outs Change Everything Traditional SEO trained you to optimize for one query, one intent, one piece of content. The mechanics of LLMs are different. When a user asks a generative engine a complex question, the model breaks it down before answering. A question like "what type of company is Vurto right for" might generate fan-outs like: - "Vurto GEO tool features" - "GEO for e-commerce 2026" - "LLM monitoring for SMBs France" - "best tool to track ChatGPT citations" - "AI readability e-commerce site" The AI's final response is built from sources found across each of these sub-angles. If you're only present on the primary query but missing from the derivatives, you fall out of the final answer — partially or entirely. ## What SEO Teams Do Instead They optimize for keywords. Keywords that humans typed into Google. Those keywords don't map to LLM fan-out queries. That doesn't mean SEO work is useless — it builds a base of content that feeds LLMs. But stopping there means missing the specific logic of generative engines. LLM fan-out queries have different characteristics. They tend to be longer. They focus on specific attributes, comparisons, usage conditions. They vary by LLM: ChatGPT and Perplexity don't decompose the same queries the same way. The result: a content strategy calibrated purely on classical search intent is partly disconnected from what actually drives visibility in AI responses. ## How Fan-Outs Work LLMs now have web search tools they activate based on the nature of the query. The exact architecture varies by model and evolves fast. But the decomposition mechanism is structural. The model identifies information gaps in its knowledge base (what we call grounding queries — more on that in another article). It then builds a series of sub-queries to fill those gaps. Then it synthesizes the results into a coherent response. What matters for your GEO visibility is being present on the right derivatives — not just on the initial question. ## The Content Strategy That Addresses Fan-Outs Understanding the fan-outs in your category means identifying the sub-angles to build or strengthen content around. Map the fan-outs for your strategic queries. Ask LLMs questions about the key topics in your category and observe how they break down their searches. Some tools let you visualize these decompositions directly — Vurto does this, identifying the fan-out queries associated with your brand and sector to calibrate content recommendations. Produce content dedicated to specific attributes. Fan-outs often target precise dimensions: use cases, comparisons, access conditions, measurable results. A piece of content that directly answers "what's the ROI of a GEO strategy for an e-tailer" is more likely to be pulled than a generic GEO article. Cover the multi-LLM spectrum. ChatGPT and Perplexity don't generate the same fan-outs. Perplexity leans toward journalistic sources and recent comparisons. ChatGPT favors well-structured content with strong expertise signals. The optimal content strategy covers both logics. Be present where LLMs look. Publishers, LinkedIn, Reddit, YouTube — these platforms are increasingly mobilized as sources in fan-outs. Being present only on your own domain means playing on a single field. ## What This Changes for Content Teams The brief changes. Instead of "an article on GEO for SMBs," the right brief becomes "a piece of content answering the question 'is GEO relevant for a company with under 50 employees and no dedicated SEO team' with concrete data and examples." The difference is subtle. The impact on AI visibility is not. LLMs pull content that answers a sub-question directly and precisely. Informational density counts. Clarity of the answer counts. Format counts. ## The Signal Most Teams Don't Have Knowing which fan-outs touch your category is a tangible competitive advantage. Not because it's an esoteric technique, but because it's a signal the vast majority of competitors don't have. In 2026, most marketing teams are still optimizing for Google keywords. That's useful. It's not enough. Brands that add fan-out logic to their content strategy are the ones starting to appear in AI answers where their competitors are absent. Today's blind spot is tomorrow's lever. --- *Vurto identifies the fan-out queries associated with your brand and sector to calibrate GEO content recommendations.* ### E-commerce ### [GEO and E-Commerce: Your Catalog as an AI Visibility Asset](https://vurto.ai/en/resources/geo-and-e-commerce-your-catalog-as-an-ai-visibility-asset) (2026-06-21) E-tailers have a GEO advantage nobody tells them about. They already have the content. Thousands of product pages with descriptions, specs, use cases. Informational material that LLMs love to draw from when answering their users' purchase questions. The problem: this material is in the wrong format. It was built for Google and for human conversion, not for LLMs. And the gap between the two is wider than most people think. ## What LLMs Do with Purchase Queries "What vacuum cleaner for a 1,000 sq ft home with a dog?" "Which earbuds for working out under $80?" "Best mattress for side sleepers?" These queries land in ChatGPT, Perplexity, and Gemini in massive volume. Users want recommendations, not lists of results. And LLMs respond with recommendations — often very specific, often with justifications based on product characteristics. Who gets cited in these recommendations? Rarely the merchants' product pages. Usually comparison sites, buying guides, review articles. Because that content is structured to answer a question, not to sell a product. The e-tailer loses the direct recommendation and sees their traffic arrive from a third-party comparison site that talks about their product — when they're lucky. ## Why Classic Product Pages Miss LLMs A typical product page is built like this: image carousel, price, prominent buy button, spec bullet points, short description, dynamically loaded customer reviews, similar product section. For a human on the site, it works. For an LLM scanning the page for information, it's a nearly empty page. The carousel doesn't load. The reviews are in JavaScript. The description is short and generic. The spec bullet points exist but explain nothing about the use case. Result: the LLM can't build a rich recommendation from this page. It looks elsewhere. ## The Markdown Conversion: What It Changes Converting a product catalog to markdown transforms a format designed for visual rendering into one designed for text processing. In markdown, the product page changes in nature. The title becomes a precisely named entity with its key attributes. The description becomes a prose paragraph that explains who the product suits and why. Technical specs sit in a cleanly structured table. Use cases are described in natural language. Comparisons with alternatives are stated clearly. This version isn't visible on the site. It's accessible to LLM crawlers via a dedicated URL or through the llms.txt. It doesn't replace the classic product page — it complements it for a different audience. The impact on AI visibility can be significant. An LLM that finds a well-structured markdown version of a product page has exactly the material it needs to formulate a precise, sourced recommendation. ## The Ten-Criteria Audit Before rewriting everything, you need to know where the gaps are. An audit against ten LLM readability criteria gives a precise picture of the catalog's state and the right priorities. These criteria cover: title precision, quality of prose description, presence of explicit use cases, mobilizable numerical data, named entities, review accessibility, technical spec structure, comparative positioning, alignment with llms.txt, and content readability without JavaScript. Vurto offers this audit across an entire catalog, with per-page scoring and prioritization of which SKUs to tackle first based on commercial importance and LLM visibility potential. The tool then rewrites the pages in real time in markdown according to the identified standards, with validation of the result against each of the ten criteria. ## The Product Categories Where GEO Changes the Most Not all categories face LLMs equally. Some are massively touched by purchase queries directed at generative engines. Consumer electronics. Home appliances. Tools. Beauty and skincare. Sports equipment. Books and board games. Any category where users need to compare before buying and where technical criteria matter — that's where LLMs get queried. In these categories, not being cited in AI recommendations means not existing in a growing share of the purchase journey. Conversely, in impulse-buy or commodity categories, the impact is less direct. The person buying socks or printer paper isn't asking Perplexity for advice. ## The Right Way to Start Not by rewriting the entire catalog at once. Start with the 50 to 100 most commercially important SKUs and check their LLM readability score. Often 20% of SKUs drive 80% of traffic. Those are the ones to tackle first. Then audit the site's technical structure to verify that these pages — even well rewritten in markdown — are actually accessible to LLM crawlers. Perfect content blocked by robots.txt is perfectly useless content. GEO for e-commerce isn't a revolution. It's a logical extension of existing SEO and content work. But it requires shifting the reference frame: thinking "LLM readability" alongside "human conversion," and producing formats that serve both. --- *Vurto analyzes product pages against ten criteria and converts them to structured markdown to maximize their presence in LLM recommendations.* ### [Product Pages and AI: Why Your Catalog Is Unreadable](https://vurto.ai/en/resources/product-pages-and-ai-why-your-catalog-is-unreadable) (2026-06-16) A product page was designed for two readers: the human consumer and Google's bot. The human looks at the photo, the price, the reviews, the name. Google's bot indexes the title, the meta, the Schema.org structured data. Thirty years of e-commerce refined this alchemy. The product pages of major online retailers are conversion and ranking machines. They're often unreadable to LLMs. ## What LLMs Read — and What They Ignore Language models process text. Not JavaScript. Not image carousels. Not tabs hidden behind lazy loading. Not descriptions buried under a "read more" that requires a click. The typical e-commerce product page structure looks like this: an H1 title, a few bullet points of specs, a condensed description block, visuals, a price, a customer reviews section. For an LLM scanning the page, the useful signal is often buried in HTML noise. The direct consequence: when a user asks ChatGPT or Perplexity "what's the best [product] for [use case]," the model pulls sources that discuss those products in a dense and structured way — comparisons, expert reviews, buying guides. Rarely the merchant's official product page. The e-tailer loses the recommendation. Sometimes to a comparison site or media outlet that talks about their product better than they do. ## The Ten Criteria That Make a Product Page Readable for AI There's nothing mysterious about it. But it requires an honest audit. **1. Descriptive, precise title.** The title must contain key attributes: brand, model, main feature, target use case. A vague title like "Men's sports shoe" is useless for an LLM trying to answer "what running shoe for a beginner trail runner under $100." **2. Structured prose description.** Bullet points listing technical specs aren't enough. A written description explaining who the product suits, in what context, and why — that's what the LLM can work with. **3. Explicit use cases.** "Ideal for..." isn't a use case. "Designed for 2-to-5-day mountain hikes with a load under 33 lbs" is a use case the LLM can actually mobilize. **4. Honest comparisons.** Product pages carefully avoid comparing themselves to competitors. LLMs love comparisons. A page that clearly positions the product against alternatives — without disparagement — gives models the material to answer comparative questions. **5. Measurable data.** Weight, dimensions, battery life, yield, capacity, warranty length. LLMs pull numerical data in comparative queries. A purely qualitative description sends the model looking elsewhere — often to a competitor. **6. Properly named entities.** Brand, technology, material, certification, partner — these entities must appear explicitly in the text, not just in metadata. **7. Content accessible without JavaScript execution.** If the main description loads in JS, the LLM crawler doesn't see it. Testing a product page with a text-only crawler routinely surfaces surprises. **8. No generic duplicated content.** Descriptions copied from the supplier or mass-generated without differentiation read as noise. The LLM prefers a source with a distinct angle. **9. Reviews integrated and accessible.** Customer reviews often contain the most useful phrasing for an LLM ("perfect for wide feet," "holds well on descents," "great for beginners"). If they load dynamically and are invisible to crawlers, that's lost signal. **10. Links to contextual content.** An isolated product page is a weak signal. A product page linked to a buying guide, a comparison, a blog post — that's a node in a content network that LLMs value more. ## The Specific Challenge of E-Commerce Catalogs The problem retailers have with LLMs isn't just one badly structured page. It's scale. A catalog with 5,000 SKUs means 5,000 product pages to reassess through the LLM lens. Manual rewriting is impossible. Generic AI rewriting, without calibration to LLM readability standards, produces text that looks readable to a human but remains poorly structured for generative engines. This is exactly what Vurto addresses for e-tailers: a real-time catalog rewrite in markdown format, structured according to LLM readability criteria, with each product page scored against ten dimensions before and after rewriting. The output isn't a generic version of the pages — it's a version calibrated so that LLMs can process, understand, and recommend them. ## What This Changes for Sales Measuring the direct impact of an LLM-ready product page on revenue is tricky, because traffic from AI engines is still partially opaque. But the signals are there. Direct referrals from ChatGPT and Perplexity are now trackable via UTMs and referrers. Brands watching this traffic see that it converts well — the user arrives with a qualified intent, not mere curiosity. They've already received a recommendation from the AI before landing on the page. Being the cited reference in that recommendation means capturing high-intent traffic that traditional advertising doesn't reach. ## The Five-Minute Test Open ChatGPT. Ask: "[your flagship product name]: does it work for [your customers' main use case]?" Watch the response. If the model cites your product page or your domain — good. If it cites an external comparison, a press article, or worse, a competitor — you know what needs fixing. --- *Vurto analyzes product pages against ten AI readability criteria and rewrites them in structured markdown to maximize their presence in LLM recommendations.* ### GEO Technical ### [Schema.org in 2026: Which Structured Data Still Matters](https://vurto.ai/en/resources/schema-org-in-2026-which-structured-data-still-matters) (2026-08-24) Ahrefs tested the theory on 1,885 pages in 2026. Add schema.org markup, measure the effect on citations in AI Overviews. Result: no improvement. Citations even dropped 4.6%. On the other side, 71% of pages cited by ChatGPT contain structured data, and Google itself confirms that schema helps models understand a page's content. Two serious studies, two conclusions that contradict each other. This is exactly the kind of confusion that pushes a marketing team either to dump JSON-LD everywhere as a precaution, or to drop the topic altogether, assuming it no longer matters now that classic SEO has changed nature. Both reactions are mistakes. ## The myth that refuses to die The misunderstanding comes from confusing correlation with cause. Pages cited by LLMs often have schema, that is true. But they also often have a clear content structure, direct answers, and established domain authority. Schema accompanies those qualities; it does not produce them. A product page stuffed with microdata on a slow site, with vague copy and zero third-party proof, will not get cited any more than before. Google says it in black and white in its documentation: no special files, no specific schema.org markup required to appear in AI features. That is not an invitation to ignore the topic. It is a reminder that markup is not a shortcut to visibility, contrary to what part of the GEO industry claims when it sells schema as a magic formula. The Ahrefs test confirms this reading without invalidating it entirely. Across 1,885 pages, adding generic, poorly targeted schema on content that already failed to answer a precise question changed nothing. That is consistent. An LLM does not cite a page because it wears a JSON-LD label. It cites a page because it contains the clearest, best-formulated, most verifiable answer to the question asked. Schema helps the model extract that answer faster and with less ambiguity. It does not create the answer. ## What actually works, and what is useless Not all schema types are equal in front of generative engines. FAQPage comes first, and by a wide margin, because it structures content into self-contained question-answer pairs, exactly the format an LLM can extract and cite without rewriting. Combining FAQPage with Article and HowTo produces nearly twice as many citations as Article alone. The logic is simple: the closer the markup brings the page structure to the format of a conversational answer, the more it eases the model's work. | Schema type | Real usefulness for GEO in 2026 | |---|---| | FAQPage | Strong. Question-answer format directly usable by LLMs | | HowTo | Strong, especially combined with Article or FAQPage | | Product | Useful for e-commerce, verifiable price and availability | | Organization | Useful for entity clarity — who you are, what you do | | Review / AggregateRating | Useful as third-party proof, provided it is authentic | | BreadcrumbList | Weak direct impact on citations | | Generic Article without structured content | Almost none | The table reads in a precise direction. Markup that describes a content structure already designed to answer a question has a real effect. Markup that merely decorates an existing page without changing its editorial structure has a marginal effect, or none. It is the difference between putting a label on a dish that is already well prepared, and sticking a label on an empty plate hoping it will fill itself. Product schema deserves a separate mention for e-commerce. Price, availability, verified reviews: this structured data lets an LLM answer a transactional question without hallucinating an outdated price. A product page without Product schema forces the model to guess or to cite a third-party source, often a comparison site that has no obligation to put you forward. ## The invisible infrastructure nobody sees running The real role of schema in 2026 is not direct citation. It is entity clarity. An LLM that encounters your site must quickly understand who you are, what you sell, where you operate, and what your relationship is with the entities mentioned in your content. Organization schema, with the right fields filled in (legal name, industry, service area, links to social profiles and the knowledge base), reduces ambiguity on that point. It guarantees no citation. It does, however, prevent a model from confusing your brand with a namesake, or from wrongly attributing to you information that belongs to a competitor with a similar name. This infrastructure function explains why Microsoft and Google confirm that schema helps models understand content, without ever promising that it triggers a citation. Schema builds the context in which your content is interpreted. It never replaces the quality of that content. A company that thinks of GEO only in terms of technical markup makes the classic mistake of confusing the scaffolding with the building. There is a second, more discreet layer to this. LLMs build internal knowledge graphs from what they crawl. A site with consistent markup across all its pages, entities correctly linked to each other, a clear hierarchy between organization, products and editorial content, becomes easier to integrate into that graph. A site where each page uses different, incomplete, sometimes contradictory markup makes that integration harder. The result is not visible in an immediate citation metric. It shows over time, in the stability and precision of what models know how to say about you. ## Prioritize without drowning in markup The question to ask is never "how many schema types can I add" but "which pages have content structured enough for the markup to be useful". A poorly written FAQ page, with vague answers, will gain nothing from carrying FAQPage markup. You first have to rewrite the answers so they are self-contained, verifiable, directly citable, then add the markup that makes them easy to extract. The priority order that makes sense in 2026 comes down to three moves. First, Organization schema on institutional pages, to settle once and for all who you are in the eyes of the models. Then, FAQPage and HowTo on content that already answers precise, frequent questions — the ones your fan-out queries reveal. Finally, Product schema on pages that carry a price and availability, if you are in e-commerce. Everything else — BreadcrumbList, decorative markup on generic content pages — can wait or be ignored without measurable consequence. Schema.org in 2026 is no longer a magic lever, and it never really was. It is an accelerator of understanding for content that already deserves to be understood. On weak content, it catches nothing up. On strong content, it removes friction between your page and the answer a model will formulate from it. *Vurto scores your product pages and key pages against ten AI-readability criteria, markup included, to spot where schema actually helps and where it is only decorating a page that needs something else.* ### [LinkedIn and GEO: Why LLMs Love This Platform (and How to Use It)](https://vurto.ai/en/resources/linkedin-and-geo-why-llms-love-this-platform-and-how-to-use-it) (2026-06-23) ## What LLMs Look for in a Source A large language model, when generating a response, doesn't browse the web the way a human does. It draws on two things: its training memory (the billions of documents ingested before its cutoff date) and, for models with web access, a real-time search layer that retrieves and synthesizes fresh sources. In both cases, a source's perceived quality depends on specific signals: - Semantic density: does the content say something precise, or is it filler? - Factual consistency with other sources: if 12 documents converge on the same information, the model treats it as reliable. - Clear attribution: who is speaking, in what capacity, on what subject. - Readable structure: clean HTML, explicit headings, continuous prose. LinkedIn meets these criteria better than almost any social network. A well-written LinkedIn post cites a figure, names a company, identifies an author with a title and experience. That's exactly what an LLM wants to anchor a factual response. ## Why LinkedIn Survives Training Better Large language models were trained on Common Crawl, Reddit dumps, Wikipedia, books, and press articles. LinkedIn is part of these corpora, in various forms depending on licensing agreements. What matters is that LinkedIn is a platform built around professional identity. Every piece of content is attached to a profile: a name, a job title, a company, a sector. This richness of context makes each post a micro-document dense with attribution signals. "Marie Dupont, Marketing Director at X, writes about B2B content strategy" is more useful to an LLM than an anonymous article published on a dateless, authorless blog. LinkedIn articles — the long-form feature, distinct from short posts — have a stable URL, a structured title, content that tends to be longer and better referenced. They resemble blog posts, but with a LinkedIn identity layered on top. For a model trying to attribute an opinion or expertise, that's ideal. The profile itself is a source. When an LLM answers "who is the French reference on marketing automation," it can synthesize information from LinkedIn profiles, posts, and articles published on the platform. Your LinkedIn presence isn't just a digital CV. It's a node in the knowledge graph that LLMs draw from. ## What Makes LinkedIn Content Citable Not all LinkedIn posts are equal from an LLM's perspective. A post with five emojis, a clickbait hook, and a "tell me in the comments" has little chance of being cited by a model answering a serious question. What LLMs retain is high-information-density content. A few concrete criteria: **Sourced figures.** A post that says "our conversion rate increased 34% after restructuring our landing pages according to these three principles" is citable. A post that says "here are 5 tips to convert better" isn't really — too generic, too diluted. **Clear positions.** LLMs synthesize expert opinions. If your content asserts something specific, on a given subject, with your name and title attached, it becomes a potential source for answering "what do experts think about [topic]?" **Exact industry vocabulary.** Models recognize technical language. A post on GEO that uses the right terms (fan-out queries, grounding, llms.txt, share of voice) will be better understood and better semantically indexed than a post that paraphrases the same concepts in plain language. **Consistency.** A profile that publishes weekly within a coherent thematic perimeter builds a strong semantic footprint. The model recognizes it as a specialized source on that subject. A profile that publishes on ten different topics generates noise. **Useful length.** Short posts (under 150 words) perform poorly unless they contain a standout fact. Long-form articles are more easily ingested as documentary sources. The ideal GEO format is probably the mid-length post (300–600 words) with clear structure and actionable information. ## The LinkedIn Strategy for Corporate GEO For a brand, LinkedIn plays a dual role in GEO strategy. First role: amplifying the presence of your experts. When your directors, consultants, and product managers publish on LinkedIn using your brand name and product vocabulary, they create distributed occurrences that strengthen your signal in LLMs. One company blog versus ten active profiles: the ten profiles win on volume and signal diversity. Second role: occupying comparison queries. LLMs often answer questions like "what's the best solution for [problem]" by synthesizing expert opinions and experience reports. LinkedIn is a natural source for this type of content. If your clients publish positive feedback on LinkedIn mentioning your product by name, those posts can end up in the LLM synthesis of a comparison question. The distinction matters: this isn't about publishing promotional posts. LLMs don't cite disguised ads. They cite expert opinions grounded in concrete experience, data, and precise examples. ## What LinkedIn Can't Do Alone LinkedIn is a powerful lever, but in isolation it's not enough. LLMs build their responses through triangulation. If your brand appears on LinkedIn but nowhere else — no press coverage, no in-depth articles on your site, no mentions in specialized forums like Reddit or industry communities — the signal is fragile. The most robust GEO strategy combines multiple surfaces: your site (with content readable by AI crawlers), your LinkedIn presence, mentions in third-party publications, and ideally a presence in community discussion spaces where LLMs look for authentic opinions. LinkedIn is a central piece of the setup, not the whole setup. What changes compared to classic SEO is that each surface now has direct value for your LLM visibility, independently of backlinks. An excellent LinkedIn post with no backlinks can be cited by ChatGPT. In traditional SEO, it would have carried little weight. ## What You Should Do This Week Audit your LinkedIn presence through the GEO lens. Ask yourself three simple questions. Does your profile describe your expertise clearly, using the exact vocabulary of your sector? LLMs read titles, summaries, and experience sections. If your profile uses vague language ("helping companies navigate their transformation"), it won't be cited as a reference on a specific topic. Do your recent posts contain citable factual information? Figures, examples, clear positions on a subject. Or is it surface content designed to generate likes? Is your company mentioned in your colleagues' and clients' posts in a substantive way, with precise details about what you do and for whom? If the answer is no on all three points, you have a straightforward, fast opportunity to improve your GEO signal without touching your site or launching a complex content campaign. --- *Vurto monitors your presence in ChatGPT, Gemini, Perplexity, and Claude in real time. If you want to know whether your brand or your LinkedIn experts are being cited by LLMs today — and in what context — that's exactly what Vurto measures.* ### [Your Site Is Invisible to AI — and You Don't Know It](https://vurto.ai/en/resources/your-site-is-invisible-to-ai-and-you-don-t-know-it) (2026-06-20) There's a difference between a site Google can index and a site an LLM can read. Google sends a crawler that retrieves HTML, follows links, logs tags. It's a well-documented mechanism. Thousands of tools audit it. Millions of developers optimize for it. LLMs work differently. When a model goes looking for information on your site, it doesn't navigate. It scans. It extracts text. It looks for structured meaning. And if what it finds is opaque, fragmented, buried in JavaScript or poorly tagged HTML, it moves on. The result: sites that are perfectly optimized for Google — well-ranked, fast-loading, green Core Web Vitals scores — that are nearly nonexistent from an LLM's perspective. ## What "Readable for AI" Actually Means It's not a question of technical performance. It's not meta-descriptions or alt tags either. AI readability is the capacity of a page's content to be extracted, understood, and mobilized by a language model when answering a user's question. It starts with a simple question: if you strip all the JavaScript from your page and keep only the raw HTML, what readable text is left? On many modern e-commerce or SaaS sites, the answer is: not much. The main content is rendered dynamically. Descriptions hide behind accordions. Customer reviews come in via a third-party API. Prices appear after a JavaScript call. For an LLM, none of that exists. ## The Most Common Blind Spots The first blind spot is JavaScript-rendered content. An LLM scanning a product page and finding only an empty HTML shell — because all the content is rendered client-side — will simply ignore the page. Modern frontend frameworks are great for user experience. For LLM readability, they're often a disaster. The second is invisible structure. Dense content without clear hierarchy, without well-positioned headings, without a readable progression — an LLM struggles to use it. It doesn't read like a human. It looks for structural signals to identify what matters. The third is the absence of markdown. LLMs process markdown natively. Well-structured markdown content is more readily usable than equivalent content in rich HTML. For strategic pages — product pages, comparison pages, buying guides — having a markdown-accessible version is a direct advantage. The fourth is a poorly calibrated robots.txt. Blocking certain bots without realizing they're precisely the crawlers used by LLMs. GPTBot, ClaudeBot, Googlebot AI Mode, PerplexityBot — agents that can end up blocked by robots.txt rules copy-pasted from templates that predate their existence. ## The AI Readability Audit in Practice The audit starts with a simple test: access each of your strategic pages via a text-only browser or a JavaScript-free scraper and observe what's left. It's usually revealing. Then check the coherence between robots.txt and your GEO strategy. If certain crawlers are blocked, is it intentional? If so, which LLMs are affected? Is the impact on visibility accepted? Check the status of the llms.txt. Does it exist? Is it consistent with the site's actual content? Does it point to the right pages? Analyze the informational density of key pages. A product page, a comparison page, a pricing page — do they have enough textual substance for an LLM to extract something useful? Vurto offers a dedicated tool for this diagnosis: a site-wide AI readability check, page by page, that identifies technical blockers, content invisible to crawlers, robots.txt/llms.txt inconsistencies, and pages to prioritize for rewriting or restructuring. It's often the first tool web teams open when they start working on GEO. ## What This Changes for Technical Teams GEO isn't just a marketing topic. It has direct technical implications. Server-side rendering or partial server-side hydration for strategic content. Static markdown files for product pages or guides. A robots.txt audit to verify LLM crawler access. llms.txt implementation. Up-to-date and coherent Schema.org structured data. None of these actions are complicated. But they require coordination between SEO, tech, and content teams that doesn't yet exist in most organizations. The problem isn't technical. The problem is that nobody has clearly asked the question yet: are our strategic pages readable by LLMs? And in the vast majority of cases, the honest answer would be: not really. --- *Vurto checks your site's AI readability: content invisible to crawlers, robots.txt, llms.txt, and the structure of strategic pages.* ### [llms.txt: The robots.txt of the AI Era](https://vurto.ai/en/resources/llms-txt-the-robots-txt-of-the-ai-era) (2026-06-17) In 1994, site owners started placing a `robots.txt` file at the root of their domains. It told web crawlers what to index and what to ignore. Nobody forced them. It became a universal standard because it solved a real problem: giving bots the instructions they needed to do their job properly. Thirty years later, a new file is emerging with the same logic: `llms.txt`. Its purpose? Give language models the essential information about your site, structured in a way they can process efficiently — without having to piece it together themselves. ## What llms.txt Does The `llms.txt` file lives at the root of a domain, just like `robots.txt`. It contains a structured description of the site — in markdown, readable by LLMs — with key information about the entity, its products, its values, its trusted sources, and its strategic pages. It can include things like: - What the company does, in one precise sentence - Main products or services and their characteristics - Target audience - The most important pages to know about - Authors or experts associated with the site - Links to third-party sources that mention the site (press, studies, comparisons) The principle: rather than letting an LLM reconstruct an incomplete or distorted image of your brand from scattered fragments, you give it a structured, verifiable, coherent information base to work from. ## Why It Matters for GEO LLMs have two ways of knowing a brand. The first is training. Everything ingested before the model's knowledge cutoff. That's where well-known brands live — established facts, information repeated across many sources. For recent, niche, or poorly represented brands, this layer is partial or inaccurate. The second is real-time search. When the model has access to search tools, it goes looking for current information. And what it finds first are the best-structured, most easily parseable sources. `llms.txt` plays on both levels. It improves the quality of information available to the model in real time, and it structures that information in a way that maximizes the likelihood the model uses it in its responses. ## What a Poorly Written llms.txt Does (or Doesn't Do) An empty or generic llms.txt is a missed opportunity. An auto-generated llms.txt without calibration is sometimes worse — it can contain vague descriptions that reinforce an inaccurate image. The common mistakes: Description too vague. "Innovative company in the digital sector" tells an LLM nothing. "GEO platform that tracks brand citations in ChatGPT, Gemini, and Perplexity and recommends content actions based on fan-out queries" says a great deal. No named entities. LLMs operate on entities — names, products, technologies, locations, people. An llms.txt without clear entities is barely usable. Missing third-party sources. Pointing to press articles, comparisons, and studies mentioning your brand reinforces the credibility of the information in the model's eyes. Misalignment with the rest of the site's content. If the llms.txt says one thing and the site's pages say another, the model arbitrates based on the strongest signals — not necessarily in your favor. Contradictory robots.txt. Blocking AI crawlers in robots.txt while maintaining a llms.txt is incoherent. Both files need to be aligned. ## llms.txt and robots.txt: Coherence as a Condition The llms.txt doesn't work in isolation. It connects with robots.txt, with Schema.org structured data, with the overall quality of the site's content. The coherence between robots.txt and llms.txt is a point often neglected. Some sites block AI crawlers in robots.txt by instinct — data protection reflex — without realizing it completely neutralizes the llms.txt. Others allow crawling but have an llms.txt that contradicts their product pages. Vurto includes a dedicated module for auditing and rewriting both files together — robots.txt and llms.txt — to ensure the instructions given to models are coherent, complete, and aligned with the overall GEO strategy. It's one of the starting points of any serious GEO audit. ## How to Write an Effective llms.txt The recommended structure is simple: ```markdown # [Company Name] [Precise description in 2-3 sentences: what you do, for whom, with what difference] ## Products / Services - [Product 1]: [functional description in one sentence] - [Product 2]: [functional description in one sentence] ## Target Audience [Description of the primary audience: sector, size, role, problem solved] ## Key Pages - [URL page 1]: [what's there] - [URL page 2]: [what's there] ## Reference Sources - [Press article link] - [Comparison link] - [Study citing the brand] ``` This isn't a fixed format — the standard is evolving fast. But the underlying logic is stable: precise information, named entities, links to third-party sources, coherence with the rest of the site. ## Where the Market Stands in 2026 llms.txt adoption is growing. Among the most GEO-advanced sites, it's become a baseline reflex, just like robots.txt or a sitemap. For brands that haven't done it yet, it's one of the GEO actions with the best effort-to-impact ratio. A few hours of work, a well-written file, a robots.txt coherence check — and you give LLMs a reliable information base about your brand that they didn't have before. In a world where generative engines build their responses from whatever signals are available, controlling those signals at the source is basic digital hygiene. --- *Vurto audits and rewrites llms.txt and robots.txt files to ensure coherence and maximize your site's readability for LLMs.* ### Brand & Reputation ### [Brand Perception in AI: What ChatGPT Really Says About You](https://vurto.ai/en/resources/brand-perception-in-ai-what-chatgpt-really-says-about-you) (2026-06-18) Ask ChatGPT, Perplexity, and Gemini this question today: "[Your brand name]: what are the pros and cons?" The answers will surprise you. Sometimes in a good way. Often not. What LLMs say about your brand is the result of an aggregation of sources: press articles, customer reviews, forums, comparisons, product pages, LinkedIn posts, industry studies. The model synthesizes all of it into a view of your brand — and that view can diverge significantly from your official positioning. That's AI brand perception. And it's a signal almost nobody is tracking right now. ## Why AI Brand Perception Is Different from Classic Reputation Management Online reputation management has been around for twenty years. Google Alerts, social mentions, platform reviews. Mature tools for a well-understood problem. Brand perception in LLMs is different for several reasons. LLMs aggregate and synthesize. They don't show you a list of mentions — they give you a consolidated view. And in that consolidation, some signals carry more weight than others, following an opaque logic. LLMs have partial memory. What a model knows about your brand depends on its training cutoff, the sources ingested during training, and what it finds in real-time searches. A negative event from two years ago can still weigh heavily if the sources covering it are well-indexed. A recent positive story may be absent if it hasn't been crawled yet. LLMs aren't consistent with each other. ChatGPT, Gemini, Perplexity, and Claude don't share the same sources or the same response architectures. Your brand perception can be positive in one and neutral or negative in another. AI brand perception influences purchase decisions. That's what makes this strategically important. When a prospect asks an AI about your brand before buying — and that's a rapidly growing behavior — what they receive shapes their decision. Not in the same way as a Trustpilot review, but potentially in a more synthetic, more authoritative way. ## The Dimensions of AI Brand Perception Perception isn't a single dimension. It breaks down into several layers. Overall sentiment. Do AIs talk about your brand in positive, neutral, or negative terms? This is the most visible layer and the easiest to communicate to leadership. Associated attributes. When LLMs describe your brand, what adjectives and qualifiers do they use? "Reliable," "premium," "accessible," "innovative," "hard to onboard" — these attributes build an image that influences prospects. Attributed use cases. Who do AIs recommend your brand to? A tech startup? An enterprise? A beginner? LLMs often have fixed ideas about the use cases of brands, sometimes out of step with your actual commercial reality. Comparisons. When you're compared to a competitor in an AI response, do you come out ahead, behind, or even? And on which criteria? Sources mobilized. Which articles, comparisons, and sites do LLMs cite when they talk about your brand? These sources shape perception and may contain outdated or inaccurate information. ## The Common Gaps Between Real Image and AI Image In practice, the most frequent gaps involve: Target positioning. A B2B mid-market brand may be described as an SMB solution because the comparisons mentioning it target smaller businesses. The AI copies the bias of its sources. Features. Recent features launched after the model's cutoff or poorly documented publicly are absent from the description. The brand is presented with its capabilities from eighteen months ago. Pricing. LLMs often cite outdated price ranges pulled from old comparisons. For brands whose pricing has evolved, this directly affects conversion. Friction points. Old negative feedback on a specific issue — customer service, a product transition period, a bug fixed long ago — can continue to weigh on AI perception if the sources covering it remain well-indexed. ## How to Correct AI Brand Perception The fix isn't magic. It takes work on the sources. Identify which sources carry the most weight. Which articles and comparisons do LLMs most often cite when covering your brand? Those are the sources to prioritize: update them, complement them, or counterbalance them with new ones. Produce content that challenges incorrect attributes. If AIs associate your brand with a use case you no longer target, create explicit content about your current audience — on your own site and with third-party publishers. Models update their perception as they crawl. Work on the density of positive named entities. LLMs pull named entities to build their responses. Content that explicitly associates your brand with positive attributes ("fast deployment," "measurable ROI in 30 days," "works for teams without dedicated SEO resources") feeds perception in the right direction. Monitor continuously, not once. AI brand perception evolves. A minimum monthly check is needed to catch drift before it becomes entrenched. Vurto integrates this dimension directly into its dashboard: brand sentiment measurement across LLMs, identification of associated attributes, tracking of mobilized sources, and alerts when perception diverges from the defined positioning. It's the monitoring tool that classic brand reputation management doesn't offer yet. ## The Signal No Marketing Team Should Ignore What AIs say about your brand is what millions of users receive as their first answer about you. Not an ad. Not a sales pitch. A synthesis that the model presents as factual information. Ignoring it means letting others shape your image — your past sources, your old reviews, your better-AI-referenced competitors — without ever taking back control. The good news: this perception is correctable. It takes time, content, and method. But it's not set in stone. --- *Vurto measures brand perception across LLMs: sentiment, associated attributes, mobilized sources, competitor comparisons.* ### Monitoring & Analytics ### [Measuring GEO ROI: The Question That Blocks Every Budget](https://vurto.ai/en/resources/measuring-geo-roi-the-question-that-blocks-every-budget) (2026-09-21) 70.6%. That is the share of traffic generated by AI chatbots that arrives on a site with no origin trace at all. No "referred from ChatGPT", no "referred from Perplexity". Nothing. Google Analytics dumps it all into the "Direct" bucket, the same one used for a visitor who typed the site address from memory. A GEO budget (optimizing visibility in chatbot answers from ChatGPT, Claude, or Gemini, as opposed to classic Google search ranking) can therefore produce real results and still remain invisible in the dashboard used to justify it. That is the heart of the problem. A marketing team invests in GEO, gets citations in AI answers, sometimes sees traffic rise, but cannot answer the simplest question a CFO will ask: how much does it return. Not because GEO returns nothing. Because measurement tools were built for a world where people clicked blue links, not for a world where they get a direct answer inside a conversation window. ## Why the click is no longer enough Return on investment (the calculation that compares what an action cost to what it returned) has always relied on a chain of trackable clicks. A user clicks a Google result, lands on a page, fills a form or buys a product. Each step leaves a trace, each trace feeds a dashboard. That chain breaks with AI chatbots, for three precise technical reasons. The mobile apps for ChatGPT or Claude do not pass origin information when they open a link. Many users copy-paste an address straight from the chatbot answer instead of clicking, which erases any provenance trail. And ChatGPT Plus, like Google's AI mode, deliberately adds a technical instruction that prevents destination sites from knowing where the visitor came from. As a result, a site can see its "Direct" traffic double without anyone on the team understanding why. The answer is almost always the same: part of that supposedly originless traffic actually comes from AI chatbots. And this ghost traffic is not anecdotal. It converts at 10.2%, versus 2.46% for non-AI traffic classified elsewhere. A visitor from a chatbot buys or fills a form four times more often than an average visitor, but nobody knows because nobody sees it. Before trying to calculate ROI, the first task is therefore to repair measurement itself. A channel grouping that manually recognizes known AI domains (chatgpt.com, perplexity.ai, claude.ai and the like) recovers part of the identifiable traffic. For the rest, the traffic that arrives with no trace at all, only server-log analysis, more tedious but more reliable than classic analytics tools, can rebuild a faithful picture. ## The metrics that replace the click Once measurement is fixed, a deeper problem remains. Even when well measured, the click tells only a small part of the GEO story. A majority of Google searches now end with no click at all, and that share rises sharply whenever an AI-generated summary appears at the top of the results page. A user who asks ChatGPT a question and gets a satisfying answer, citing the brand along the way, often has no reason to click anywhere. The brand was still seen, named, recommended. That value does not land in any clicks column. Substitution metrics are then needed, ones that measure presence rather than passage. The first is citation frequency: on a sample of queries representative of the sector, in how many answers does the brand appear, and at what rank in the source list. The second is share of voice, which compares that frequency to competitors on the same queries. The third is the sentiment attached to the citation: an answer that mentions the brand in positive terms does not weigh the same as a neutral mention in the middle of a list, or worse, a mention paired with a caveat about quality or price. The fourth is the nature of the sources the chatbot goes looking for to build its answer, because a brand absent from the pages the AI consults to form an opinion has no chance of being cited, whatever its budget. These four metrics do not replace the financial calculation. They build the base without which that calculation means nothing. A GEO budget that moves citation frequency and share of voice forward, month after month, produces an effect that will eventually translate into traffic and sales, even if the exact path between the two stays blurry. Conversely, a budget that moves none of these four numbers will never produce a return, no matter how good the published content is. ## What a chatbot visitor is really worth The most counterintuitive point in this whole topic is that AI traffic, when measured correctly, converts clearly better than classic traffic. Several studies published between late 2025 and 2026 agree on an order of magnitude: between four and five times the average conversion rate of classic organic search, with gaps ranging from 1.3x on impulsive online purchases up to more than twenty times in some B2B sectors where the buying decision is long and deliberate. | Traffic origin | Average conversion rate | |---|---| | Classic Google search | 1.76% | | Claude | 5.0% | | Perplexity | 10.5% | | ChatGPT | 15.9% | The logic behind these numbers is simple once you see it. A user who clicks a Google result is still comparing, opens several tabs, hesitates. A user who already queried a chatbot has gotten a synthesis, asked follow-up questions, eliminated some options. When they finally click through to a site, it is often because they are ready to act. The chatbot did part of the qualification work upstream. This data changes how a GEO budget must be read. A lower volume of AI traffic than classic Google volume is not necessarily a failure, if that rarer traffic converts four times better. The right question is not "how many visitors did GEO bring", it is "how many sales or leads did this traffic, even reduced, produce". It is a complete reversal of how classic search ranking learned to think for twenty years, when volume was king. ## Building a formula that holds up Without waiting for a perfect tool that does not exist yet, a company can build a reasonable GEO ROI measure on four pillars. The first is a starting snapshot, taken before any investment: current citation level, share of voice versus competitors, dominant sentiment. Without that baseline, there is no way to know whether the efforts produce an effect. The second is regular tracking of those same visibility metrics, month after month, to objectify progress rather than relying on a feeling. The third is reconstituting traffic that actually came from chatbots, via a correctly configured channel grouping and, when possible, server-log analysis to recover what escapes standard tools. The fourth is linking that reconstituted traffic to the conversions it produces, keeping in mind that its conversion rate will likely be higher than other channels, not lower. Once these four building blocks are in place, the calculation becomes: the value of sales or leads attributable to that traffic, minus what GEO cost, all relative to the starting investment. It is not exact science. It is a considerable improvement over the total absence of measurement, which remains the norm today in most companies that are already investing in this topic. The temptation, faced with this complexity, is to wait for analytics tools to catch up before acting. That is the inverse error of spending without ever measuring. Companies that start today tracking their citation frequency and share of voice are building a history. Those that wait will start in a year with no baseline at all, at the exact moment their competitors already have twelve months of it. *This is exactly what Vurto does: track citation frequency, share of voice, and brand sentiment across ChatGPT, Gemini, Perplexity, and Claude, month after month, to give marketing teams the measurement base that most GEO budgets still lack.* ### [AI Chatbot Traffic Is Tiny. It's Worth Gold](https://vurto.ai/en/resources/ai-chatbot-traffic-is-tiny-its-worth-gold) (2026-09-10) A US agency, Seer Interactive, compared two sources of visitors on the same site: those arriving from ChatGPT, and those arriving from a classic Google search. Result: 15.9% of the former end up buying or filling out a form, versus 1.76% of the latter. Almost nine times more. Microsoft, using its Clarity analytics tool, observed the same thing across more than 1,200 media and publisher sites: visitors coming from a chatbot sign up eleven times more often than those coming from a classic Google search. A study of 94 e-commerce sites measured a more modest but real gap: 31% more conversions for ChatGPT traffic compared to a Google search that doesn't include the brand name. These numbers vary from study to study, sometimes a lot, because the calculation methods differ. But they all point in the same direction. Traffic sent by an AI chatbot converts noticeably better than classic traffic. The problem is that it accounts for barely more than 1% of a site's total traffic, according to measurements from the analytics tool Conductor. A marketing team that only looks at volume misses the most useful piece of information on its dashboard. ## Why a visitor from a chatbot buys more often The answer lies in how the question was asked before the person even arrived on the site. When someone types "best invoicing software for freelancers" into Google, they get ten links and still have to compare, sort, and hesitate. When they ask the same question to ChatGPT, the model has already done that work. It has read several sources, compared the options, and returns only the two or three names that best answer that person's specific request. The visitor who clicks the final link is no longer searching, they're verifying a decision that's already been nearly made. It's the difference between a customer who walks into a store to browse, and a customer who walks in already knowing what they want to buy, on the recommendation of a friend they trust. The chatbot plays that role of trusted intermediary. It has filtered, explained, compared. The person landing on your site no longer needs to be convinced of the product's general value, only reassured on the details. A report published in March 2026 confirms the scale of the phenomenon: visitors coming from AI assistants like ChatGPT or Perplexity convert 42% better than non-AI traffic, across all channels combined. The quality of the intent that precedes the visit changes the nature of the visitor itself. Take a concrete example. Someone looking for payroll software for ten employees types their question into Perplexity. The engine answers by citing three tools, each with an indicative price, a strength, and a limitation. They click on the one that best fits their situation, land on a pricing page, and request a demo that same day. The same buyer, starting from a classic Google search, would probably have opened five or six tabs, compared reviews for an hour, then put off the decision until the next day. The chatbot didn't just shorten the path. It did the filtering on their behalf, and that filtering translates directly into conversion rate. ## Advertising is still useful, but for a different goal This isn't about pitting advertising against visibility in AI chatbots as two competing strategies. They serve different needs. Advertising pays for immediate, controllable volume: you know how much you're spending, how many visitors you're getting, and you can adjust the budget from one day to the next. Getting cited by an AI chatbot doesn't work the same way. It's built over time, through clear content, visible customer reviews, technical documentation that the models reading it can access. The conversion gap between the two channels doesn't mean you should cut your advertising budget. It means that a euro invested in the quality of what AI chatbots find about a brand produces a return that shows up in no advertising campaign dashboard, and that return will grow mechanically as buyers increasingly use these tools. ## What this changes in how you measure performance Most marketing dashboards are built around volume: number of visits, number of ad clicks, number of impressions. That logic makes sense for paid advertising, where each visit has a direct, predictable cost. It becomes misleading for traffic generated by AI chatbots, where volume will stay modest for a long time but where each visit carries more weight. The table below summarizes the gap observed across several recent studies, with different methodologies, which explains the variation in the numbers: | Traffic source | Observed conversion rate | Relative volume | |---|---|---| | Classic Google search | Baseline (1x) | High | | Visitors from an AI chatbot | 4x to 11x depending on the study | Low, often under 1% of the total | | Paid advertising | Variable, heavily dependent on cost per click | High but costly to sustain | Comparing a channel that accounts for 1% of traffic to a channel that accounts for 60% in raw volume doesn't make sense. The right question isn't "how many visitors does the chatbot send us," it's "what is the value of a visitor sent by a chatbot, compared to that of a visitor from an ad." On that basis, a budget devoted to being well cited in AI chatbot answers can pay off more, visitor for visitor, than an equivalent ad budget, even if the total number of visitors remains incomparable for now. ## A signal that will carry increasing weight Traffic from AI chatbots remains marginal today, but its trajectory is clear: it grows every quarter, while classic organic traffic flattens out. A team that invests now in its visibility within the answers given by ChatGPT, Perplexity, or Gemini isn't just building one more channel. It's capturing a stream of pre-qualified visitors, fewer in number but far more likely to become customers, before this channel gets big enough for everyone to chase it at once. The logic of classic advertising is still useful for generating volume quickly. But a budget aimed entirely at volume, with nothing invested in the quality of what arrives via generative engines, ignores the underlying trend in buying behavior. The next quarter won't be won just by attracting more people, but by attracting the right people, at the right moment in their decision. *Vurto measures the share of your traffic that genuinely comes from citations earned in ChatGPT, Perplexity, Gemini, and Claude, to tell apart what's small in volume but big in value.* ### [The GEO Metrics That Actually Matter](https://vurto.ai/en/resources/the-geo-metrics-that-actually-matter) (2026-06-22) SEO has its metrics. Average position, click-through rate, impressions, organic traffic. Years of practice have built a shared reference frame. Everyone knows what a good DA means. Everyone reads the same dashboards. GEO doesn't have that yet. The market is young. Tools are being built. And in that fog, many GEO dashboards measure what's easy to measure — not what counts. Here are the signals worth tracking. ## Citation Rate, Not Raw Presence The first obvious metric is presence: is my brand mentioned in LLM responses? That's level zero. Useful for detecting total absence. Not enough to steer by. What counts is the weighted citation rate: across prompts related to my category, in what proportion of responses do I appear? And above all, is that proportion moving over time? A citation rate that grows is a sign that content, activation, and monitoring work is paying off. A stable rate despite content investment signals either more active competition or a technical readability problem neutralizing the effort. ## Sentiment by LLM Being cited is good. Being cited positively is the goal. Brand sentiment in LLMs is a metric that didn't exist two years ago. The most advanced GEO tools now measure it: for each LLM, what's the overall tone of brand mentions? Positive, neutral, negative? And which specific attributes drive that sentiment? The multi-LLM dimension matters here. Positive sentiment on Perplexity and negative on ChatGPT is a common situation. It usually signals different sources being mobilized by each model — an old comparison visible to ChatGPT, recent positive press accessible to Perplexity. Identifying that is knowing where to act. ## Share of Voice by Prompt Category A competitor can have a lower overall citation rate than you and still beat you on the prompts that convert the most. Segmentation by prompt category is essential: comparison prompts ("X vs Y"), purchase intent prompts ("best X for Y"), discovery prompts ("what is X"), validation prompts ("is X reliable"). Across each of these angles, share of voice can vary dramatically. A B2B brand might be well-positioned on discovery prompts and absent from comparison prompts — where purchase decisions happen. That's a precise diagnostic that guides content priorities. ## Quality of Sources Mobilized LLMs don't cite all the content that shaped their knowledge. They cite a subset of sources — typically the most recent, best-structured, or highest-trust-signal ones. Tracking which sources LLMs mobilize when they talk about your brand is a powerful diagnostic metric. If it's consistently third-party comparisons you don't control, dating back two years, containing outdated information — you know where to work. If your own pages start appearing as sources, that's a signal of technical progress (AI readability, llms.txt) and editorial progress. ## Trend Over Time, Not Snapshots All these metrics are nearly useless as point-in-time readings. GEO is a medium-term practice. A citation rate moving from 12% to 18% over six months on purchase prompts in your category is a strong signal. A snapshot at 18% without a baseline is noise. Trend measurement requires consistency. Weekly for sentiment and citation rate on key prompts. Monthly for competitive share of voice. Quarterly for the audit of mobilized sources. Vurto structures this monitoring continuously, with granularity by LLM, by prompt category, and by competitor. The dashboard gives access to the trend, not just the current number — which is rare among market tools, most of which offer static views. ## What Shouldn't Be in a GEO Dashboard A few metrics that show up frequently and shouldn't be steering a GEO strategy. Raw total mention count. Without weighting for context quality or prompt relevance, this number means little. A brand can be mentioned 500 times a week in negative or off-target contexts. Rankings in "top GEO tools" lists. Some platforms build their own prompt bases biased toward benchmarks that favor them. Watch out for self-referential rankings. Presence on obscure LLMs with near-zero market share. Monitoring ten LLMs of which eight have 0.1% usage share is noise. Focusing on ChatGPT, Gemini, Perplexity, and Claude covers the vast majority of real generative traffic. ## The Question to Ask Every Week Is my LLM visibility growing on the prompts that match my actual buyers? Not "am I mentioned," but "am I mentioned in the right conversations, with the right framing, from the right sources?" That's the question that turns GEO monitoring into strategic management. --- *Vurto monitors citation rate, brand sentiment, competitive share of voice, and mobilized sources by LLM — continuously.* ## Academy (learning guides — full content) Full Vurto Academy curriculum with the full content of each lesson, module by module (74 lessons). ### Technical setup ### [Why everything starts with the technical layer](https://vurto.ai/en/academy/setup-technique/pourquoi-tout-commence-par-la-technique) (2026-05-17) Without a solid technical foundation, content cannot be crawled, indexed, or cited by a generative AI — the technical layer is the entry condition for any visibility strategy, not a step you handle "later." It governs a three-link chain — crawl, indexing or ingestion, then citation — and if the first link breaks, the other two never get a chance to happen. ## The chain: crawl → indexing/ingestion → citation Before a page can show up in Google results or in a ChatGPT answer, it first has to be found. That's the job of **crawling**: a bot (Googlebot, Bingbot, or an AI crawler like GPTBot) follows links across the web, reads a page's HTML, and decides whether to keep it for later use. Next comes **indexing** on the classic search engine side: the page gets added to a database that lets it surface for relevant queries. On the generative AI side, the equivalent is usually called **ingestion**: content gets folded into document stores or indexes the model can consult when generating an answer — through real-time web search for Perplexity or browsing-enabled ChatGPT, or through datasets used at different stages depending on the provider. The last link, **citation**, depends entirely on the first two. A generative AI cannot cite a page it was never able to read, and it won't read it if a bot could never reach it in the first place. It's a simple point that gets overlooked constantly: teams optimize content with citation in mind while the actual blocker sits one or two steps earlier, at the level of technical access. ## What breaks this chain in practice In most technical audits, the same blockers show up again and again: - An overly restrictive **robots.txt** that blocks entire sections of a site, sometimes by mistake or out of excessive caution. - Client-side JavaScript rendering with no server-side rendering or pre-rendering, which some crawlers fail to execute properly. - **Orphan pages** with no internal links pointing to them, invisible to a bot that navigates link by link. - Server errors (5xx), redirect loops, or slow response times that push bots to reduce how often they come back. - No up-to-date **XML sitemap**, leaving search engines without a clear list of pages to explore. - Duplicate content or a confusing URL structure, which spreads crawl budget thin across pages with no real value. Each of these looks minor in isolation. Combined, they explain why some sites with genuinely solid content stay invisible. ## SEO and GEO share the same foundation GEO doesn't replace technical SEO — it builds on top of it. AI crawlers operate through mechanisms fairly close to those of classic search engines: they read HTML, generally respect the robots.txt file, and are sensitive to whether a page is actually reachable. It's generally observed that a page invisible to Googlebot also has little chance of being properly read by an AI crawler, since the same technical obstacles — unrendered JavaScript, server errors, robots.txt blocks — tend to affect both kinds of bots in similar ways. That's why a GEO strategy that ignores the technical layer starts at a structural disadvantage. You can produce the best-structured content in the world — if it isn't reachable, it stays invisible both to someone typing a Google query and to someone asking Claude or Perplexity a question. ## A quick test to see where you stand Before diving into tools, a few immediate checks give you a first diagnosis without installing anything: - Type `site:yourdomain.com` into Google: if the number of results is far lower than your site's actual page count, a chunk of your content is likely not indexed. - Open `yourdomain.com/robots.txt` in a browser: if a `Disallow: /` line appears under `User-agent: *`, your entire site is closed off to bots, including AI crawlers. - Turn off JavaScript in your browser (or use Search Console's "Inspect URL" tool once it's set up) and reload a key page: if the main content disappears, a bot that doesn't execute JavaScript properly will see the same thing. - Check whether a sitemap exists at `yourdomain.com/sitemap.xml`: its absence isn't a dealbreaker on its own, but it deprives search engines of a useful reference point. These four checks take less than five minutes and are often enough to tell whether a visibility problem comes from an obvious technical block before digging any further. ## Where to actually start Before working on substance — content structure, direct answers, structured data — you need to secure the basics. In order: 1. Check what's already indexed today (a `site:yourdomain.com` search on Google, or a direct look at Search Console if it's already set up). 2. Install and configure Google Search Console for a reliable view of indexing status. 3. Audit the robots.txt file and XML sitemap to make sure nothing important is being excluded. 4. Check loading speed and Core Web Vitals, which influence crawl budget. 5. Only after that, work on structuring content to make it "GEO-ready." This lesson sets the general frame. The next ones get concrete: installing [Google Search Console](/en/academy/setup-technique/installer-google-search-console), understanding the [robots.txt file and controlling AI crawlers](/en/academy/setup-technique/robots-txt-et-crawlers-ia), and digging into [why AI sometimes ignores your site](/en/academy/comprendre-le-geo/pourquoi-les-ia-ignorent-votre-site) once the technical basics are in place. ### [Installing and configuring Google Search Console](https://vurto.ai/en/academy/setup-technique/installer-google-search-console) (2026-05-18) Setting up Google Search Console takes three steps: create a property (domain or URL-prefix), prove you own it, then submit your XML sitemap. The whole process takes about fifteen minutes and is the first building block of any technical tracking, for SEO as much as for GEO. ## Create an account and choose the right property type Go to `search.google.com/search-console` with a Google account and click "Add property." You'll be offered two options: - **Domain property**: covers the entire domain, across all subdomains and protocols (http, https, with or without www). This is the recommended option in most cases, since it gives you a complete view and avoids duplicating properties. - **URL-prefix property**: covers only one exact URL (for example, only `https://www.yourdomain.com`). Useful if you can't edit your domain's DNS records, or if you want to track a specific subfolder in isolation. If you have control over your domain's DNS configuration, go with the domain property — it simplifies everything else downstream. ## Verifying the property: the available methods Google needs to confirm you actually control the site before granting access to its data. Several methods exist depending on the property type you chose: 1. **DNS record (TXT)** — the only method available for a domain property. Google gives you a value to add as a TXT record with your host or DNS registrar. Propagation can take anywhere from a few minutes to 48 hours. 2. **HTML meta tag** — for a URL-prefix property, Google provides a `` tag to insert in the `
` of the homepage. 3. **HTML file upload** — a uniquely named file (e.g. `google1234567890.html`) provided by Google, placed at the site root and reachable directly at `yourdomain.com/google1234567890.html`. 4. **Google Analytics or Google Tag Manager** — if either is already installed and you have admin rights on it, verification happens automatically with no extra step. Pick whichever method matches your technical access: DNS if you manage the domain name, meta tag or HTML file if you only have access to the site's code. ## Submitting the sitemap Once the property is verified, go to the "Sitemaps" entry in the side menu. Enter your XML sitemap URL (usually `sitemap.xml` or `sitemap_index.xml`, appended after your domain) and click "Submit." This step isn't required for Google to crawl your site, but it speeds up discovery of new pages and gives Google an explicit list of the URLs you consider important. If you're not sure whether a sitemap already exists, just test `yourdomain.com/sitemap.xml` directly in your browser. Most CMSs (WordPress with Yoast or Rank Math, Shopify, Webflow) generate one automatically. ## Extra settings worth doing right away A handful of settings, done once the property is verified, save time down the line: - **Link Google Analytics 4** if you already have it installed, so you can later cross-reference crawl and traffic data. - **Add other team members as users** (Settings > Users and permissions menu), with an access level that matches their role (read-only or full owner). - **Check the "Configuration" > "Changes"** section, which lists structural changes Google has detected (title tag changes, canonical changes, etc.) — useful for spotting unintended modifications. - If the site is multilingual, don't expect Search Console to handle international targeting automatically: that depends on your hreflang implementation, which needs to be checked separately. Once this setup is done, the next step is learning to [read your data in Google Search Console](/en/academy/setup-technique/lire-donnees-google-search-console) to turn this access into an actionable diagnosis. It's also worth [installing Bing Webmaster Tools](/en/academy/setup-technique/installer-bing-webmaster-tools), which follows the same logic and matters for visibility in Copilot's answers. ### [Reading your Google Search Console data](https://vurto.ai/en/academy/setup-technique/lire-donnees-google-search-console) (2026-05-19) Four Google Search Console reports deserve regular attention: Performance (what actually drives clicks), Indexing (what's really indexed), Sitemaps (what Google managed to crawl), and Core Web Vitals (loading experience). Together, they give you a complete diagnosis without needing any third-party tools. ## The Performance report: queries, pages, clicks, impressions This is the most-viewed report, and often the most poorly read one. It crosses four metrics: **impressions** (how many times a page showed up in results), **clicks**, **CTR** (click-through rate, clicks divided by impressions), and **average position**. A few useful readings: - A page with a lot of impressions but a low CTR usually points to a weak title tag or meta description, not a content problem. - An average position stuck around page 2 (positions 11-20) signals a page Google considers relevant enough to exist, but that's missing the signal needed to break into page 1 — often internal linking or content depth. - The query filter lets you spot keywords you're already ranking for without realizing it, sometimes unexpectedly — a solid starting point for complementary content. Always segment by device type (mobile vs. desktop) and by country if your audience is international: global averages often hide significant gaps. ## Coverage / Indexing: what's indexed and what isn't The "Pages" report (formerly "Coverage") lists your URLs by status: indexed, or excluded with a specific reason. The most common reasons worth watching: - **"Crawled – currently not indexed"**: Google visited the page but chose not to index it, usually due to low perceived value or duplicate content. - **"Discovered – currently not indexed"**: Google knows the URL exists but hasn't crawled it yet, generally a sign of limited crawl budget or low perceived priority. - **"Page with redirect"** or **"404 error"**: fix it or leave it as-is, depending on whether it's intentional or an oversight. - **"Duplicate, Google chose different canonical than user"**: Google picked a different canonical URL than the one you declared, often a sign of poorly managed duplicate content. The goal isn't 100% of pages indexed at any cost: some pages (filters, technical pagination, internal search results) have no reason to be. The goal is for the pages that matter to be indexed, and to understand why the ones that aren't, aren't. ## Sitemaps: checking everything is accounted for In the Sitemaps report, Google shows the number of URLs discovered in the submitted sitemap. A large gap between the number of URLs in the sitemap and the number of pages actually indexed (visible in the Pages report) is a warning sign — it can point to a poorly generated sitemap, outdated URLs, or low-quality pages Google refuses to index despite their presence in the sitemap. ## Core Web Vitals in Search Console The Core Web Vitals report aggregates real user data (from the Chrome User Experience Report) and buckets your pages into three categories — good, needs improvement, poor — across three metrics: LCP (loading time of the largest visible element), INP (responsiveness to interactions), and CLS (visual stability). Unlike a one-off testing tool, this report reflects actual visitor experience, grouped by URL pattern, which helps with prioritization: it's more valuable to fix a page template used by 500 URLs than to fix one isolated page. ## Secondary reports worth not ignoring Beyond the four main reports, two others deserve a regular look even though they get checked less often: - **The Links report** lists your most frequent external backlinks, your most internally linked pages, and the anchor text used. A central page on your site that gets very few internal links is a sign it risks being under-prioritized by bots, regardless of how good its content is. - **The "Security Issues" report** flags any detected hacking or misleading content identified by Google. A compromised site can see its crawl frequency drop sharply, an easy thing to miss if this report never gets checked for lack of a visible alert elsewhere. These two reports are rarely the priority on a first pass, but ignoring a security issue for several months can undo all the effort put into every other lever. ## Cross-referencing the data to prioritize The most effective method is cross-referencing reports rather than reading them in isolation: 1. Spot high-potential pages in Performance (lots of impressions, average position 8-20). 2. Check in Pages whether they're fully indexed without restriction. 3. Check in Core Web Vitals whether their template is being held back by a speed issue. 4. Prioritize fixes on the templates that affect the most high-potential pages at once. This cross-reading is exactly the foundation of a full technical audit: the next lesson details how to [audit your site with generative AI and GSC data](/en/academy/setup-technique/auditer-son-site-avec-ia-et-gsc) to speed up this prioritization, before moving on to the [complete technical checklist](/en/academy/bases-seo-techniques/checklist-technique-complete). ### [Auditing your site with generative AI + GSC data](https://vurto.ai/en/academy/setup-technique/auditer-son-site-avec-ia-et-gsc) (2026-05-20) Exporting your Google Search Console data as CSV and feeding it to ChatGPT or Claude with a structured prompt gets you a prioritized list of technical fixes within minutes — as long as you cross-reference several exports and verify every recommendation before applying it. ## Why run GSC data through a generative AI Google Search Console shows raw data: tables of queries, pages, indexing statuses. Nothing in it automatically tells you what to fix first. A generative AI is useful here not because it "knows" your site, but because it can quickly cross-reference several columns of data, spot recurring patterns (say, a page template that keeps showing up among non-indexed URLs), and turn a technical diagnosis into plain language for a team that isn't made up of SEO specialists. This isn't a replacement for human analysis, it's an accelerator: what would take an hour of manual sorting in a spreadsheet can be done in a few minutes, provided you feed it clean data and a precise prompt. ## Exporting the right data from Search Console Before involving an AI, prepare your exports. Three reports are enough for a first-pass audit: 1. **Pages report** (Indexing): export the full list of URLs with their status and exclusion reason. This is where most technical blocking signals live. 2. **Performance report**: export data by page (not just by query), over a period of at least 3 months, with impressions, clicks, CTR, and average position. 3. **Core Web Vitals report**: export the breakdown by URL group if volume allows. Each export is done via the "Export" button in the top-right corner of each report, as CSV or Google Sheets. Clean up the files if needed (empty columns, duplicates) before handing them over. ## Building an effective audit prompt A vague prompt gets a vague result. Structure your request with context, objective, and expected format. For example: > "Here's an export of my Pages report from Google Search Console for [domain name], a [type of business] site. Analyze the exclusion statuses and identify which reasons come up most often. For each recurring reason, explain the likely cause and propose a concrete fix. Rank the results by number of affected URLs, from most to least impactful." For the Performance report, a useful prompt is to ask it to spot high-potential pages: "Identify pages with more than 500 impressions and an average position between 8 and 20 — these are candidates for targeted optimization rather than a full rewrite." Avoid generic prompts like "how do I improve my SEO": without attached data and format constraints, the answer stays theoretical and hard to act on. ## What an AI is good at spotting — and what it isn't A generative AI excels at: spotting statistical patterns across a large number of rows, grouping URLs by template or site section, and translating a technical report into plain language for a marketing team. It's more limited when it comes to: verifying in real time whether a fix has actually been applied (it has no access to your live site unless you explicitly provide it), judging the real severity of an issue without business context (a poorly indexed page can be entirely normal if it holds little value), and replacing an actual technical test (it can't run PageSpeed Insights on your behalf, for instance). Every recommendation it generates should therefore be verified before implementation, especially anything touching robots.txt or redirects, where a mistake can deindex important pages. ## Turning the analysis into a prioritized action plan Once you have the analysis, always ask for a usable format: a table with columns "issue," "number of affected URLs," "proposed fix," "estimated effort." That table, not the raw text of the answer, is what becomes your technical roadmap. This method builds directly on the report-reading approach from the previous lesson on [reading your Google Search Console data](/en/academy/setup-technique/lire-donnees-google-search-console). Once technical fixes are underway, it's worth checking that your [site speed and Core Web Vitals](/en/academy/setup-technique/vitesse-du-site-core-web-vitals) are keeping up, then digging into [how a search engine reads a page](/en/academy/bases-seo-techniques/comment-un-moteur-lit-une-page) to understand the fundamentals behind these diagnostics. ### [Installing Bing Webmaster Tools](https://vurto.ai/en/academy/setup-technique/installer-bing-webmaster-tools) (2026-05-21) Bing Webmaster Tools takes a few minutes to set up, with a direct import available straight from Google Search Console — it's the quickest step in this module, and one of the most overlooked, even though it directly affects your visibility inside Copilot. ## Why Bing matters for GEO It's natural to associate Bing with a modest market share compared to Google in classic search. But Bing plays a different role in the generative AI ecosystem: Copilot, Microsoft's assistant, relies in part on Bing's index for its web-search-backed answers. A site that's absent or poorly indexed on Bing has fewer chances of being picked up by that assistant, regardless of its standing on Google. That alone is reason enough not to skip this tool, even if your current Bing traffic looks marginal. ## Create an account and add your site Go to `bing.com/webmasters` and sign in with a Microsoft account (or create one). Click "Add a site" and enter your domain's URL. ## Import directly from Google Search Console If your site is already verified in Google Search Console, Bing Webmaster Tools offers an import option: you authorize access to your Google account, and Bing automatically pulls in the list of your verified properties, your already-submitted sitemap, and a sample of your performance data. This is by far the fastest method — it skips a new property verification and a new sitemap export. ## Verifying ownership if you don't import If you'd rather verify manually, or your site isn't on Google Search Console yet, three methods are offered: 1. **HTML meta tag** to insert in the `` of the homepage. 2. **XML file upload** at the site root (similar to Google's HTML file, but in XML format). 3. **DNS record (CNAME)** to add with your registrar, for domain-level verification. Pick whichever matches your level of access, exactly as with Google Search Console. ## Exploring Bing-specific reports Once the property is active, Bing Webmaster Tools offers reports similar to Google Search Console's (Performance, Crawl, Site Index) but also a few tools of its own: - **SEO Reports**, which automatically lists technical issues detected on your pages (missing titles, duplicate meta tags, images without an alt attribute), with no extra setup required. - **URL Inspection**, the equivalent of Google's URL inspection, which lets you check how Bingbot sees a specific page and request an on-demand crawl. - **Backlinks**, a report listing external links pointing to your site, useful alongside Google Search Console's Links report for cross-referencing both sources. These reports need less upkeep than Google's, but are worth a monthly check once the tool is set up, particularly the SEO Reports section, which often surfaces simple technical fixes. ## Submitting the sitemap and setting the basics Once the property is verified, go to the "Sitemaps" menu and submit your XML sitemap URL, just like with Google. Then check two useful settings: - **Targeted country/region**, in the site settings, if your business is local or national. - **Crawl control**, a Bing-specific feature that lets you throttle how often the bot visits if your server has limited resources — only touch this if you're seeing a real impact on server performance. Setting up Bing Webmaster Tools rounds out the technical tracking foundation started with [Google Search Console](/en/academy/setup-technique/installer-google-search-console). With both tools in place, the next step is adding [Google Analytics 4](/en/academy/setup-technique/installer-google-analytics-4) to measure, this time, the traffic actually generated by visitors coming from generative AI. ### [Installing Google Analytics 4](https://vurto.ai/en/academy/setup-technique/installer-google-analytics-4) (2026-05-22) Installing Google Analytics 4 properly takes four steps: create the property and data stream, add the tracking code, wire up consent management, then verify data is actually coming in before you rely on it to analyze traffic from generative AI. This installation builds on the foundation already laid with [Google Search Console](/en/academy/setup-technique/installer-google-search-console) and [Bing Webmaster Tools](/en/academy/setup-technique/installer-bing-webmaster-tools): it adds traffic measurement on top of crawl measurement. ## Create a GA4 property and data stream On `analytics.google.com`, create a property (if you're starting from scratch, GA4 is now the only version Google offers). Fill in the timezone and currency, which can't easily be changed later. Then add a **data stream** of type "Web" with your site's URL: this stream generates a measurement ID (formatted `G-XXXXXXX`) used to connect your site to this property. ## Install the tracking code Two main methods: 1. **Direct gtag.js tag**: paste the code snippet Google provides into the `` of every page on the site. Simple for a small static site, but rigid if you plan to add other tracking tools later. 2. **Google Tag Manager (GTM)**: recommended as soon as the site has more than a handful of pages, or if you expect to add other tools (ad pixels, other analytics). GTM installs once via two code snippets (one in the ``, one right after the `` opens), and GA4 is then configured as a tag inside GTM — no need to touch the site's code again for every change. Most CMSs (WordPress, Shopify, Webflow) also offer native GA4 or GTM integration through a dedicated settings field, which avoids editing source code directly. ## Handle consent before any data collection If your audience includes European users, data collection through GA4 falls under GDPR and generally requires prior consent for non-essential cookies. In practice, that means connecting GA4 to a consent banner (CMP) that blocks analytics tags from firing until the user has given their approval. Google offers a native option, **Consent Mode**, which lets GA4 keep modeling some data even without consent, without storing identifying cookies. This technical point goes beyond the scope of this lesson, but it directly affects how reliable your data is: an installation that ignores consent risks skewed data or a compliance problem. ## Verify the data is actually coming in Once installed, don't assume "it works" without checking. Two tools let you confirm this immediately: - The **"Realtime"** report in GA4: open your site in a tab, browse a few pages, and check that your visit shows up within a few minutes. - The **Google Tag Assistant** browser extension or GTM's debug mode, which show exactly which tags fire, in what order, and with what data. If nothing shows up after a few minutes, the most common causes are a mistyped measurement ID, an ad blocker preventing the script from loading, or a misconfigured consent banner blocking everything by default. ## The key events to enable from the start GA4 automatically collects several events (page views, scroll, outbound clicks, file downloads) with no extra setup, as long as "Enhanced measurement" is enabled in the data stream settings — on by default, but worth double-checking. Beyond these automatic events, define at least one conversion event specific to your business (form submission, sign-up, purchase) so you can later measure whether traffic from generative AI converts differently than classic organic traffic. Once GA4 is installed and reliable, the next step is learning to [configure the GA4 reports that actually matter](/en/academy/setup-technique/configurer-ga4-rapports-utiles), particularly to isolate referral traffic from generative AI and track how it evolves over time. ### [Configuring GA4: the reports that actually matter](https://vurto.ai/en/academy/setup-technique/configurer-ga4-rapports-utiles) (2026-05-23) To track traffic from generative AI in GA4, isolate the referring domains (chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com...) in the acquisition report, ideally through a custom channel, then follow its trend over time rather than looking at a single snapshot. ## Isolating referral traffic from generative AI By default, GA4 classifies incoming traffic using predefined channel rules (Organic Search, Direct, Referral, Paid Search...). Traffic from a click on a link cited by ChatGPT or Perplexity usually lands in the **"Referral"** channel, lumped together with any other site linking to you — which makes it invisible unless you dig further. To isolate it, open the **Acquisition > Traffic acquisition** report, then add a secondary dimension for **"Session source"**. Filter on the relevant domains: `chatgpt.com`, `chat.openai.com`, `perplexity.ai`, `gemini.google.com`, `copilot.microsoft.com`, `claude.ai`. This list evolves with usage and should be reviewed regularly. ## Creating a custom channel or a dedicated segment Manually filtering every time you check gets tedious fast. Two more durable options: - **A custom channel group** (Admin > Data channel groups), where you create a rule like "if session source contains chatgpt.com OR perplexity.ai OR gemini.google.com OR copilot.microsoft.com OR claude.ai, classify as 'Generative AI' channel." This channel then shows up like any other in standard reports, allowing ongoing tracking without redoing the filter every time. - **A custom exploration** (Explore menu), letting you build a table crossing session source, number of sessions, conversion rate, and landing pages, with export available for external tracking. The custom channel is the more durable solution over time: once created, it applies retroactively to some extent and stays available across all standard reports without repeated manual work. ## Reports worth checking beyond acquisition Once the channel is isolated, three angles add more value than a simple session count: - **Landing pages**: which pages on the site actually receive this traffic? This reveals, indirectly, which content is being picked up by AI. - **Engagement rate and session duration**: does AI-driven traffic behave differently from classic organic traffic (shorter sessions, higher bounce rate, or on the contrary stronger engagement)? There's no universal rule here — it depends heavily on industry and content type, and is worth observing in your own data rather than assuming. - **Conversions by channel**: if a conversion event is set up, compare its rate on the generative AI channel against other channels, to assess the real value of this traffic rather than just its volume. ## Tracking the trend over time A one-off reading of AI traffic has little value on its own — the trend is what matters. Set up simple monthly tracking (number of sessions on the generative AI channel, compared to the previous month and to the previous year if volume allows). A gradual rise in this channel, even from a small base, is a signal that your content is starting to get picked up by answer engines — a signal that complements AI visibility scores measured elsewhere, but this time measured on real traffic rather than prompt tests. ## Setting up an automatic alert Rather than relying on remembering to check manually, GA4 lets you create **custom insights** (Admin > Custom insights) that trigger an email notification when a defined condition is met — for example, a change in session count on the "Generative AI" channel beyond a certain threshold over a week. This kind of alert removes the need to remember to check the report every month, and lets you react faster if an unusual spike or drop shows up, particularly after a robots.txt change or a content overhaul. ## Limits worth knowing This tracking has blind spots worth keeping in mind: - Some traffic from AI answers isn't necessarily attributed correctly: links shared from a mobile app or copy-pasted into a new tab can land as "Direct" traffic, invisible in the dedicated channel. - Consent banners (see the lesson on installing GA4) can reduce collected data volume if some visitors decline tracking. - Referring domain naming changes over time: a new AI engine or a domain change on the provider's side can make traffic disappear from your custom channel without warning, which is why the channel rule is worth revisiting regularly. With this tracking in place, you have a direct measure of AI traffic to weigh against the technical fixes already made through [Google Search Console](/en/academy/setup-technique/lire-donnees-google-search-console) and your [site speed](/en/academy/setup-technique/vitesse-du-site-core-web-vitals) checks. It's also a solid foundation before tackling the next module's core question: [why AI sometimes ignores your site](/en/academy/comprendre-le-geo/pourquoi-les-ia-ignorent-votre-site). ### [Site speed and Core Web Vitals](https://vurto.ai/en/academy/setup-technique/vitesse-du-site-core-web-vitals) (2026-05-24) Site speed influences both the crawl budget bots devote to your site and the actual experience of your visitors: a slow site gets crawled less often and less deeply, and loses visitors before they even reach the content. Core Web Vitals are the standard way to measure that speed in a way that's comparable across sites. ## Why speed affects crawl budget and experience **Crawl budget** refers to the number of pages a bot is willing to explore on your site within a given time. That budget isn't unlimited: it depends partly on how quickly your server responds. It's generally observed that a slow or unstable site pushes bots to reduce how often they visit and how many pages they crawl per visit, which mechanically slows down the discovery of new content and its update in the index. On the user side, the effect is more direct: a long loading time increases the abandonment rate before content even renders. For a GEO objective, this matters twice over: a slow page has less chance of being crawled by an AI crawler, and less chance of being read all the way through by a visitor who clicked through from a generated answer. ## The three Core Web Vitals metrics explained Google defined three metrics that together make up Core Web Vitals: - **LCP (Largest Contentful Paint)**: the time it takes to render the largest visible element on screen (often a hero image or a headline). Generally recommended target: under 2.5 seconds. - **INP (Interaction to Next Paint)**: the delay between a user interaction (click, keypress) and the corresponding visual update. This metric replaced FID (First Input Delay) in 2024. Generally recommended target: under 200 milliseconds. - **CLS (Cumulative Layout Shift)**: a measure of unexpected visual shifts during loading (a button that moves right before you click it, for instance). Generally recommended target: under 0.1. These thresholds are indicative benchmarks published by Google, not absolute rules that apply everywhere: a site with a lot of complex interactive content may tolerate slightly different margins depending on its use case. ## Measuring with PageSpeed Insights (and why not only that) `pagespeed.web.dev` is the reference tool for testing a specific page. It provides two types of data, worth telling apart: - **Field Data**, sourced from the Chrome User Experience Report, reflecting the real experience of your visitors over the last 28 days — only available if your site gets enough Chrome traffic. - **Lab Data**, generated from a simulated test at the moment you run the analysis — useful for diagnosis, but can vary from one run to the next depending on the simulated network conditions. The Core Web Vitals report in Google Search Console, covered in an earlier lesson, complements this tool by giving an aggregated site-wide view rather than a page-by-page one, which helps prioritize by template rather than by isolated URL. ## Mobile and desktop: two separate realities Core Web Vitals are evaluated separately for mobile and desktop, and the results often diverge quite a bit between the two. Since most web traffic today happens on mobile, and Google uses "mobile-first" indexing (your page's mobile version is the reference used for crawling and evaluation), it's advisable to treat mobile results as the priority whenever there's a trade-off to make. A site that scores well on desktop but poorly on mobile is still, from Google's standpoint, a site with a speed problem to fix. ## The most common fixes, and their order of priority In most audits, the fastest wins come from: 1. **Image compression and resizing**: serve images at the size they're actually displayed, in a modern format (WebP or AVIF), with lazy loading for off-screen images. 2. **Reducing blocking JavaScript**: defer loading of scripts that aren't essential to the initial render (live chat, third-party widgets, secondary ad pixels). 3. **Reserving space for dynamic elements**: define explicit dimensions for images, videos, and ads before they load, to avoid layout shifts (CLS). 4. **Caching and CDN**: use a content delivery network to bring static files closer to visitors, and configure appropriate browser caching. 5. **Hosting choice**: cheap shared hosting is often the hardest limiting factor to fix without switching technical solutions. ## Tracking the impact over time After each significant fix, re-test with PageSpeed Insights and watch Search Console's Core Web Vitals report over the following weeks rather than the following hour: field data is based on a rolling 28-day window, so it doesn't react instantly to a change. Once speed is under control, the site's technical foundation is largely in place. It's time to move on to controlling bot access with the lesson on [the robots.txt file and controlling AI crawlers](/en/academy/setup-technique/robots-txt-et-crawlers-ia), then to understand in detail [how a search engine reads a page](/en/academy/bases-seo-techniques/comment-un-moteur-lit-une-page). ### [The robots.txt file and controlling AI crawlers](https://vurto.ai/en/academy/setup-technique/robots-txt-et-crawlers-ia) (2026-05-25) The robots.txt file, placed at the root of your site, tells each bot which sections it can or can't crawl — including generative AI crawlers, each identified by its own user-agent (GPTBot, ClaudeBot, PerplexityBot, Google-Extended...) that you can allow or block independently of one another. ## What robots.txt does (and doesn't do) Robots.txt gives **crawling** instructions, not privacy guarantees or guaranteed indexing control. Three important clarifications: - Blocking a URL in robots.txt stops a compliant bot from crawling it, but doesn't necessarily stop it from appearing in an index if it's already known through other means (external links, for instance). To actually remove a page from an index, the `noindex` tag or a removal request through Search Console are the right tools, not robots.txt. - The file works on good faith: a legitimate, declared bot (Googlebot, GPTBot, ClaudeBot) generally respects the rules it finds there, but nothing technically stops a malicious or poorly coded bot from ignoring them. - Rules apply per **user-agent**: you can write different rules for each bot, which is exactly what lets you arbitrate between classic SEO and GEO. ## Basic syntax A robots.txt file is made of blocks, each targeting one or more user-agents: ``` User-agent: [bot name or * for all] Disallow: [blocked path] Allow: [allowed path, takes priority over a broader Disallow] ``` Simple example blocking an admin folder for all bots, but allowing everything else: ``` User-agent: * Disallow: /admin/ Allow: / Sitemap: https://yourdomain.com/sitemap.xml ``` The `Sitemap:` line at the end, while optional, is good practice: it tells bots directly where to find the full list of your pages, independent of any submission made through Search Console or Bing Webmaster Tools. ## The main AI user-agents to know Here are the most common crawlers as of today, keeping in mind this list evolves with the market: - **GPTBot** (OpenAI): used to collect content for model training. - **OAI-SearchBot** (OpenAI): dedicated to ChatGPT's web search feature, distinct from GPTBot. - **ChatGPT-User** (OpenAI): triggers when a user explicitly asks ChatGPT to look at a page in real time. - **ClaudeBot** and **anthropic-ai** (Anthropic): content collection by Anthropic. - **PerplexityBot**: crawling for indexing by Perplexity; **Perplexity-User** handles real-time browsing requests initiated by a user. - **Google-Extended**: specifically controls how your content is used by Google's AI models (Gemini), separate from Googlebot, which handles classic indexing. - **Bingbot** covers both classic Bing search and, to some extent, the data used by Copilot. Each of these user-agents can be targeted individually in robots.txt, which lets you, for example, allow Googlebot without allowing Google-Extended, or allow PerplexityBot without allowing GPTBot. ## Deciding: allow, block, or differentiate by crawler The right call depends on your objective, and there's no universal answer: - **Blocking a training crawler** (GPTBot, ClaudeBot for collection, Google-Extended) reduces use of your content for training future models, but has no effect on your presence in real-time answers generated by those same assistants, which often go through a separate web search mechanism. - **Blocking a search/citation crawler** (OAI-SearchBot, PerplexityBot, Perplexity-User) directly reduces your chances of being cited in those tools' answers, since that's the bot that lets the AI discover and read your page at the moment it answers a question. - Most sites looking to grow their GEO visibility have an interest in **allowing search/citation crawlers**, and deciding case by case on training crawlers based on their own IP and content policy. Example robots.txt that blocks training but allows citation: ``` User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: ClaudeBot Allow: / User-agent: * Allow: / Sitemap: https://yourdomain.com/sitemap.xml ``` ## The case of firewalls and bot management tools Beyond the robots.txt file itself, some sites use an application firewall or a service like Cloudflare to block or throttle bots at the network level, independent of what's declared in robots.txt. These settings are sometimes turned on by default with broad blocking rules against "bots" in general, which can block a legitimate AI crawler without any trace of it in your robots.txt. If you notice an AI crawler you've explicitly allowed never shows up in your logs, check these network-level settings in addition to the robots.txt file itself: both layers need to line up for your rules to actually produce the intended effect. ## Checking and testing your robots.txt After any change, check the file directly at `yourdomain.com/robots.txt` in a browser, and use the robots.txt tester built into Google Search Console ("Settings" > "robots.txt" report) to confirm a specific URL is allowed or blocked as intended. A syntax mistake in this file (a `Disallow: /` left under the wrong user-agent, for instance) can deindex an entire site by accident — treat it with the same care as a production change. Robots.txt works together with two other files: the [XML sitemap and the llms.txt file](/en/academy/setup-technique/llms-txt-et-sitemap-xml), which round out the setup by explicitly telling AI systems your pages and your preferences. To go further on the bots themselves, the lesson [AI crawlers: who reads you, and why](/en/academy/setup-technique/crawlers-ia-qui-vous-lit) details how to spot them directly in your server logs, beyond what they declare in their user-agent. ### [The llms.txt file and the XML sitemap](https://vurto.ai/en/academy/setup-technique/llms-txt-et-sitemap-xml) (2026-05-26) The llms.txt file is an emerging, not universally adopted standard that offers a structured Markdown summary of your site for generative AI systems; the XML sitemap, on the other hand, remains the established and widely supported standard for telling search engines the full set of pages to crawl. The two play different roles and complement each other rather than compete. ## What llms.txt is Proposed in 2024, llms.txt is a Markdown text file placed at the site root (`yourdomain.com/llms.txt`), meant to give a generative AI a condensed, readable summary of the site's content: an overview, main sections, links to the most important pages or documents. The underlying idea is to offer a compact entry point that's easier for a language model to process than a full site crawl. It isn't an official standard adopted or confirmed by the major AI providers — OpenAI, Anthropic, Google, and Perplexity haven't made any formal commitment to systematically use it. It's a community proposal, championed in particular by parts of the developer ecosystem, gaining traction without being any guarantee of adoption. ## What it delivers today (and its limits) In practice, publishing an llms.txt doesn't guarantee better citation in AI answers: no major provider has confirmed that its models or crawlers systematically consult this file when responding to a user query. What can reasonably be expected from it today: - A signal of good practice and clear structuring, useful should tools come to rely on it more broadly in the future. - A useful resource for cases where a developer or a specifically configured AI agent explicitly looks for this file (for example, certain coding assistants that use it to understand a project's documentation). - A clarifying exercise in its own right: writing an llms.txt forces you to synthesize what actually matters on the site, which has value even independent of whether an AI actually reads it. It in no way replaces the XML sitemap, robots.txt, or properly implemented Schema.org markup, which remain the best-established levers today. ## How to generate and structure an llms.txt The proposed format stays simple, in Markdown: ``` # Site name > Short one- or two-sentence description of what the site does and its value. ## Documentation - [Page name](https://yourdomain.com/page): one-line description ## Resources - [Resource name](https://yourdomain.com/resource): one-line description ``` For a content site (like a blog or a resource center), favor a list of your most foundational guides or articles rather than an exhaustive list of every page — the goal is to synthesize, not duplicate the sitemap. No universal automatic generator exists today for every CMS: writing it manually, or semi-automatically by pulling your main titles and meta descriptions, remains the most reliable method. ## A refresher on the XML sitemap and best practices The XML sitemap, meanwhile, is a standard that's been supported by every major search engine for a long time. A few best practices worth checking: - **Cap at 50,000 URLs and 50MB uncompressed per file**: beyond that, use a sitemap index referencing multiple sitemaps. - **Only include canonical, indexable, up-to-date URLs**: a sitemap containing `noindex` pages, redirects, or 404s sends a negative quality signal. - **Update it automatically** through the CMS rather than by hand, to avoid drift from what's actually published. - **Include a reliable last-modified date** (`