Insights · AI search · September 2026

How ChatGPT, Gemini and other AI engines find their answers

Your buyer asks ChatGPT for the best tools in your category, and the answer names three companies. Before you can be one of them, the engine has to find your pages. Each AI engine searches the web through an index: Google's for Gemini and AI Overviews, Bing's for Copilot, Perplexity's own, and a mix of providers for ChatGPT, Claude and Grok. Each index has a crawler you can accidentally block. A page that is public, crawlable and indexed in Google and Bing is within reach of all of them. The work that follows is the same everywhere: answer the buyer's question on a page the engines can read.

Fabian Cid

Fabian Cid

Founder, Shifter

Where each engine gets its answers

Every row below comes from the company's own documentation. These pages change often, so each row shows the date we last checked it.

Engine Where it searches The crawler that decides if you can appear Checked
ChatGPT search "Third-party search providers, as well as content provided directly by our partners" (OpenAI). For Enterprise and Edu workspaces, OpenAI says it "may share disassociated search queries with the Bing search engine" (OpenAI Help). OAI-SearchBot. "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers." GPTBot is for training (OpenAI crawlers). 28 Sep 2026
Google AI Overviews and AI Mode Google's "core Search ranking systems," which retrieve pages "from our Search index" (Google Search Central). Googlebot. The page must be indexed and eligible for a snippet (Google), and the property must be included in "Search generative AI features" in Search Console (Search Console Help). 28 Sep 2026
Gemini app "Some Gemini responses are grounded on Search results" (Gemini Privacy Hub). Deep Research uses Google Search as a source by default (Gemini Help). Google-Extended controls use of your content for Gemini training and for grounding in Gemini Apps. It "does not impact a site's inclusion in Google Search" (Google crawlers). 28 Sep 2026
Microsoft Copilot For Microsoft 365 Copilot, Microsoft says "Copilot generates a search query that it sends to the Bing search service" (Microsoft Learn). Bingbot. NOARCHIVE keeps page content out of Copilot answers; NOCACHE limits it to the URL, title and snippet (Bing Webmaster Blog). 28 Sep 2026
Perplexity Its own search index. Perplexity "partners with third-party crawlers to help build our search index" (Perplexity Help). PerplexityBot, "designed to surface and link websites in search results on Perplexity" (Perplexity crawlers). 28 Sep 2026
Claude Anthropic lists Brave Search as a web-search provider (Anthropic subprocessors). "Image search is powered by Bing" (Claude Help). Claude-SearchBot. Blocking it "may reduce your site's visibility and accuracy in user search results." ClaudeBot is for training (Anthropic). 28 Sep 2026
Grok "Internet search data is provided by internet search providers" (xAI privacy policy). In Europe, xAI names Brave (xAI Europe addendum). We found no Grok crawler or robots.txt rule in xAI's documentation. 28 Sep 2026

So there are several AI indexes, and they all crawl the same public web. That's why the checks later in this guide are short, and mostly the same for every engine.

What they cite

Knowing where an engine searches tells you whether you can appear. What it cites tells you where the competition is. Ahrefs published the most-cited domains for each engine in September 2026, based on US queries across all topics. The shares below are each domain's portion of the citations going to the top 50 domains, not of all citations.

Engine Most-cited domains (share of top-50 citations)
ChatGPT Reddit 16.8%, Wikipedia 7.0%, Consumer Reports 3.7%, Forbes 3.1% (Ahrefs)
Gemini Reddit 28.5%, YouTube 14.6%, Wikipedia 8.8%, Walmart 2.8%, Forbes 2.7% (Ahrefs)
Perplexity Reddit 21.6%, YouTube 20.8%, Wikipedia 6.3%, Facebook 4.5%, Amazon 3.6% (Ahrefs)
AI Overviews YouTube 22.9%, Reddit 18.5%, Facebook 10.1%, Google 8.8%, Instagram 5.6% (Ahrefs)
Copilot Amazon 16.8%, Walmart 12.6%, Wikipedia 7.6% (Ahrefs)

Those are all-topic figures, dominated by consumer questions. For software buyers, two studies are closer to your situation:

Treat any single percentage with care. Studies use different questions and count citations differently, and the numbers move quickly. Goodie saw Reddit's share of ChatGPT's citations fall from 8.0% to 2.9% in a single two-week stretch in mid-August 2026 (Goodie). Use studies to see patterns, then check the answers to your own buyers' questions.

What to check on your site

  1. Get indexed in Google and in Bing. Google feeds Gemini and AI Overviews. Bing feeds Copilot, and is named by OpenAI for some ChatGPT workspaces. Check Search Console and Bing Webmaster Tools.
  2. Check the "Search generative AI features" setting in Search Console. It's on by default. If it's off, your site isn't eligible for AI Overviews and AI Mode. Child properties inherit their parent's setting, so check the parent property too.
  3. Don't block the search crawlers. Leave OAI-SearchBot, Googlebot, Bingbot, PerplexityBot and Claude-SearchBot allowed. Check your CDN or firewall bot settings as well: a firewall rule can block a crawler that robots.txt allows.
  4. Decide on the training crawlers deliberately. GPTBot and ClaudeBot are for model training, and blocking them doesn't remove you from those companies' search answers. Google-Extended is different: it also governs grounding in the Gemini app.
  5. Don't hide your content from snippets. Google's nosnippet rule stops a page's content being used "as a direct input for AI Overviews and AI Mode" (Google Search Central). Bing's NOARCHIVE and NOCACHE limit what Copilot can use.
  6. Answer the buyer's question on the page. Who it's for, what it integrates with, what it costs and how it compares. Google says "You don't need to create new machine readable files, AI text files, markup, or Markdown" to appear in its AI features (Google Search Central).
  7. Get onto the pages the engines cite for your category. That usually means "best X" lists, review sites, Reddit threads and YouTube, but it varies. Look at the sources cited in answers to your own buyers' questions.

This robots.txt makes the search crawlers explicit and shows where the training crawlers go if you choose to block them:

# Search crawlers: keep these allowed to stay eligible for AI answers
User-agent: OAI-SearchBot
Allow: /

User-agent: Googlebot
Allow: /

User-agent: Bingbot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

# Training crawlers: optional. Blocking them doesn't remove you from search answers.
# User-agent: GPTBot
# Disallow: /

# User-agent: ClaudeBot
# Disallow: /

# Google-Extended also controls grounding in the Gemini app.
# Blocking it can remove you from Gemini answers.

If your file has no rules for these crawlers and no blanket Disallow: /, they're already allowed. One detail catches people out: a crawler follows only the group with the "most specific user agent" that matches it and ignores the others (Google Search Central). So once you add a group for a named crawler, it ignores your User-agent: * rules. Repeat any of those rules you still want it to follow, such as a Disallow for private folders.

Where to measure

You can't check every engine every week, so start where your buyers are. Among the seven largest AI chatbots in May 2026, ChatGPT had 53.9% of worldwide web visits, Gemini 27.9%, Claude 9.2%, and Perplexity and Copilot 1.3% each, according to Similarweb data published by Momentic (Momentic). That's why Shifter measures ChatGPT and Gemini for its clients.

Answers change from one run to the next, so one check proves little. Ask the same questions on a fixed schedule, save each answer with its date, and look at the trend. How to find buyer questions when your website has little traffic covers where the questions come from.

How we checked this

We read each engine's own documentation, and the studies cited above, between 27 and 28 September 2026. Every quote was checked against the live page, and each row of the first table shows the date of its last check. We'll recheck these pages every quarter, because four of them were updated between July and September 2026.

FAQ

Do I need separate optimization for Perplexity?

Usually not. As of September 2026, Perplexity uses its own index, but its crawler works like the others: allow PerplexityBot, keep your pages public and indexed, and answer the buyer's question clearly. Whether to measure Perplexity separately depends on whether your buyers use it. The work itself is the same.

Should I block GPTBot?

It's your decision. As of September 2026, OpenAI says GPTBot is used for training and OAI-SearchBot decides whether a site can appear in ChatGPT search answers. Blocking GPTBot doesn't remove you from ChatGPT search. Blocking OAI-SearchBot does.

Does Google-Extended affect AI Overviews?

No. As of September 2026, Google says Google-Extended doesn't affect a site's inclusion in Google Search, and AI Overviews use Google's Search index. It does control training and grounding in the Gemini app, so blocking it can remove you from Gemini's answers.

Does SEO still matter for AI answers?

Yes. As of September 2026, Google's AI Overviews and AI Mode retrieve pages from Google's Search index, Microsoft 365 Copilot sends its searches to Bing, and ChatGPT uses third-party search providers. A page that isn't crawlable and indexed can't be found by any of them.