GEO Guide

AI Search Engine List 2026: All AI Crawlers & Bots

Published July 2026 · Updated August 31, 2026 · 12 min read

AI search engines are reshaping how users find information online. But before they can cite your website, they need to crawl it. Each AI engine has its own crawler with a unique user-agent string.

This 2026 guide lists every major AI search engine, its crawler, user data, and what you need to do to make sure your site is accessible to them. We've added four new crawlers this year: DeepSeekBot, Meta-ExternalAgent, Applebot (Apple Intelligence), and cohere-ai. This 2026 guide lists every major AI search engine, its crawler, user data, and what you need to do to make sure your site is accessible to them. We've added four new crawlers this year: DeepSeekBot, Meta-ExternalAgent, Applebot (Apple Intelligence), and cohere-ai. For ready-to-copy rules, see the AI crawler robots.txt templates.

Complete AI Crawler List (2026)

AI Engine Crawler User-Agent Company Purpose
ChatGPT GPTBot OpenAI Crawls web for ChatGPT answers
Claude ClaudeBot Anthropic Crawls web for Claude responses
Perplexity PerplexityBot Perplexity AI Real-time web search for Perplexity
Google AI Overviews Google-Extended Google Feeds Google's AI features (Gemini, AI Overviews)
Common Crawl CCBot Common Crawl Open dataset used to train many AI models
Doubao (豆包) Bytespider ByteDance Crawls for ByteDance AI products
Google Gemini Googlebot Google Uses regular Googlebot (no separate crawler)
Bing Copilot Bingbot Microsoft Uses regular Bingbot (no separate crawler)
Amazon Rufus Amazonbot Amazon Crawls for Amazon's AI shopping assistant
DeepSeek DeepSeekBot DeepSeek Crawls web for DeepSeek AI search and responses
Meta AI Meta-ExternalAgent Meta Crawls for Meta AI assistant (Facebook, Instagram, WhatsApp)
Apple Intelligence Applebot Apple Powers Apple Intelligence features and Siri AI
Cohere cohere-ai Cohere Crawls for Cohere's enterprise AI models

1. GPTBot (OpenAI / ChatGPT)

GPTBot is OpenAI's web crawler. It fetches web pages to provide real-time information in ChatGPT responses. With over 300 million weekly active users in 2026, being crawlable by GPTBot is critical for AI visibility.

How to allow:

User-agent: GPTBot
Allow: /

Verify with our robots.txt AI checker

2. ClaudeBot (Anthropic / Claude)

ClaudeBot crawls the web for Anthropic's Claude AI assistant. Claude is known for its strong reasoning and coding abilities, and has grown to 50M+ monthly active users in 2026. It's widely used by developers and knowledge workers.

How to allow:

User-agent: ClaudeBot
Allow: /

3. PerplexityBot (Perplexity AI)

Perplexity is an AI-first search engine that handles 500M+ queries per month in 2026. It crawls the web in real-time to provide cited answers with source links. Unlike ChatGPT, Perplexity always shows its sources — making it especially valuable for driving referral traffic.

How to allow:

User-agent: PerplexityBot
Allow: /

4. Google-Extended (Google AI)

Google-Extended is Google's crawler specifically for AI training and AI-powered features like Google AI Overviews (formerly SGE) and Gemini. It's separate from the regular Googlebot. In 2026, Google AI Overviews appears on over 1.5 billion queries per day.

How to allow:

User-agent: Google-Extended
Allow: /

⚠️ Many sites accidentally block Google-Extended because it's a newer user-agent not in legacy robots.txt files.

5. CCBot (Common Crawl)

Common Crawl is an open-source web archive that many AI companies use to train their models. If your site is in Common Crawl's dataset, there's a higher chance AI models already "know" about you. Allowing CCBot improves your chances of being referenced in AI-generated content.

How to allow:

User-agent: CCBot
Allow: /

6. Bytespider (ByteDance / Doubao)

Bytespider is ByteDance's crawler, powering AI features in Doubao (豆包), Toutiao (今日头条), and other ByteDance products. With the Chinese AI market growing rapidly, allowing Bytespider ensures visibility in one of the world's largest AI ecosystems.

How to allow:

User-agent: Bytespider
Allow: /

7. DeepSeekBot (DeepSeek)

DeepSeekBot is the crawler for DeepSeek, one of the fastest-growing AI platforms in 2026. DeepSeek has reached 100M+ monthly active users globally, with particularly strong adoption in Asia. Its open-source models and aggressive pricing have made it a major AI search player.

How to allow:

User-agent: DeepSeekBot
Allow: /

8. Meta-ExternalAgent (Meta AI)

Meta-ExternalAgent is Meta's AI crawler, powering Meta AI across Facebook, Instagram, WhatsApp, and Threads. With Meta AI integrated into apps used by billions of people, this crawler is essential for reaching users in social media contexts. Meta AI answers are increasingly cited in WhatsApp and Facebook search.

How to allow:

User-agent: Meta-ExternalAgent
Allow: /

9. Applebot (Apple Intelligence)

Applebot powers Apple Intelligence features across iOS, macOS, and Siri. As Apple rolls out deeper AI integration across its ecosystem in 2026 — including AI-powered Spotlight search, Safari summaries, and enhanced Siri — Applebot has become a significant crawler for reaching Apple device users.

How to allow:

User-agent: Applebot
Allow: /

Note: Applebot has been around for Apple News, but is now also used for Apple Intelligence AI features. If you previously blocked it, reconsider for AI visibility.

10. cohere-ai (Cohere)

cohere-ai is Cohere's web crawler, used to feed Cohere's enterprise-focused AI models. While less consumer-facing than ChatGPT or Perplexity, Cohere powers many B2B AI applications, search tools, and enterprise chatbots. Allowing cohere-ai ensures your content is available in enterprise AI workflows.

How to allow:

User-agent: cohere-ai
Allow: /

Complete robots.txt for AI Crawlers (2026)

Here's a ready-to-use robots.txt that allows all major AI crawlers:

User-agent: *
Allow: /

# AI Crawlers — explicitly allowed
User-agent: GPTBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: CCBot
Allow: /

User-agent: Bytespider
Allow: /

User-agent: DeepSeekBot
Allow: /

User-agent: Meta-ExternalAgent
Allow: /

User-agent: Applebot
Allow: /

User-agent: cohere-ai
Allow: /

Sitemap: https://yourdomain.com/sitemap.xml

💡 Check your robots.txt now — see which AI crawlers are allowed or blocked.

AI Search Market Overview (2026)

By Users

  • • ChatGPT: 300M+ weekly active users
  • • Google AI Overviews: 1.5B+ queries/day
  • • Perplexity: 20M+ MAU
  • • DeepSeek: 100M+ MAU
  • • Claude: 50M+ MAU
  • • Meta AI: 500M+ MAU (across Meta apps)
  • • Doubao (豆包): 80M+ MAU (China)

By Query Volume

  • • Google AI Overviews: 1.5B+ queries/day
  • • ChatGPT: 200M+ queries/day
  • • Perplexity: 15M+ queries/day
  • • Bing Copilot: 80M+ queries/day
  • • DeepSeek: 30M+ queries/day
  • • Meta AI: 100M+ queries/day

How to Verify Which AI Crawlers Visit Your Site

Want to know which AI crawlers are actually visiting your site? Check your server logs. Here's how:

Option 1: Grep server logs

# Search Nginx access logs for AI crawlers
grep -iE "GPTBot|ClaudeBot|PerplexityBot|Google-Extended|CCBot|Bytespider|DeepSeekBot|Meta-ExternalAgent|Applebot|cohere-ai" /var/log/nginx/access.log

# Count visits by crawler
grep -ioE "GPTBot|ClaudeBot|PerplexityBot|Google-Extended|CCBot|Bytespider|DeepSeekBot|Meta-ExternalAgent|Applebot|cohere-ai" /var/log/nginx/access.log | sort | uniq -c | sort -rn

Option 2: Use analytics tools

Google Analytics 4 and Plausible can filter by user-agent. Cloudflare Radar and AWS CloudFront logs also show crawler activity. Look for the user-agent strings listed in the table above.

Option 3: Set up a crawler monitor

Tools like Crawlee, Datadog Bot Monitoring, or a simple cron job that parses daily logs can alert you when new AI crawlers start visiting your site. This helps you stay ahead as new AI engines emerge.

Should You Allow All AI Crawlers?

For most websites, the answer is yes. Allowing AI crawlers increases your chances of being cited in AI answers, which drives brand awareness and traffic.

You might want to restrict AI crawlers if:

  • • Your site has paywalled content (allow crawlers to preview pages, block full content)
  • • You're concerned about AI training on your content (but blocking crawlers also means losing AI search visibility)
  • • Your server can't handle extra crawl traffic (set crawl-delay instead of blocking)

The tradeoff is simple: block AI crawlers = invisible in AI search. For most businesses, the visibility gain far outweighs the risks.

FAQ

Should I use a crawl-delay for AI crawlers?

For most sites, no. AI crawlers are well-behaved and respect your server's capacity. However, if you notice server load issues from crawler traffic, you can add a crawl-delay. For example: User-agent: GPTBot
Crawl-delay: 10
. This asks the crawler to wait 10 seconds between requests. Only use this if you have evidence of server strain — an unnecessary crawl-delay may slow down how quickly your content gets indexed by AI engines.

How to block specific paths but allow others?

Use Disallow rules for specific paths while keeping the rest of your site open. For example, to block AI crawlers from your /admin/ and /private/ directories but allow everything else:

User-agent: GPTBot
Allow: /
Disallow: /admin/
Disallow: /private/

User-agent: ClaudeBot
Allow: /
Disallow: /admin/
Disallow: /private/

Repeat this pattern for each AI crawler you want to control. You can also use User-agent: * to apply path rules to all crawlers at once, then add specific Allow rules for AI bots.

Check Your AI Visibility Free

Run a free 12-dimension GEO audit — see if your site is crawlable by all major AI engines.

Start Free GEO Audit →