
Which AI bots should I allow on my website?
EasyFound · · 8 min read
Allow the bots that find and fetch pages for AI search and answers: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot and Applebot. Training bots such as GPTBot and ClaudeBot are your decision: OpenAI, Google and Apple say blocking theirs doesn't remove you from their search. On easyfound.co in late September 2026, training crawlers made about three in four of these bots' requests. Copy the robots.txt below.
You should allow the bots that fetch pages for AI search and answers. The bots that collect training data are your decision: OpenAI, Google and Apple say blocking theirs doesn't remove your site from their search.
On our own website in late September, about three in four requests from the bots below came from the two training crawlers.
Which AI bots should I allow on my website, and what does each one do?
Allow the bots that find or open pages for AI search and answers:
| Bot (company) | Job | Follows robots.txt? | Suggestion |
|---|---|---|---|
| OAI-SearchBot (OpenAI) | ChatGPT search | Yes | Allow |
| ChatGPT-User (OpenAI) | Opens a page a ChatGPT user needs | "May not apply" | Allow |
| GPTBot (OpenAI) | Training | Yes | Your choice |
| Claude-SearchBot (Anthropic) | Claude's search | Yes | Allow |
| Claude-User (Anthropic) | Opens a page a Claude user needs | Yes | Allow |
| ClaudeBot (Anthropic) | Training | Yes | Your choice |
| PerplexityBot | Perplexity's search; not training | Yes | Allow |
| Perplexity-User | Opens a page a Perplexity user needs | "Generally ignores" | Allow |
| Googlebot (Google) | Google Search, including AI Overviews and AI Mode | Yes | Allow |
| Google-Extended (Google) | A setting, not a crawler: Gemini training and Gemini app answers | n/a (a setting) | Your choice |
| Bingbot (Microsoft) | Bing search and Bing's AI chat; Microsoft may also train on it13 | Yes6 | Allow |
| Applebot (Apple) | Siri, Spotlight and Safari search7 | Yes | Allow |
| Applebot-Extended (Apple) | A setting, not a crawler: Apple's training | n/a (a setting) | Your choice |
| meta-webindexer (Meta) | Citing pages in Meta AI | Yes | Allow |
| meta-externalagent (Meta) | Training or indexing | Yes | Your choice |
Asked this question three times via OpenAI's API, ChatGPT's model never named Perplexity's, Apple's or Microsoft's bots. Yet a parent asking Perplexity for a "tuition centre in Subang Jaya with small classes" relies on Perplexity's bot reaching your page.
Does blocking AI training bots hide my business from ChatGPT, Gemini or Claude?
Blocking OpenAI's training bot doesn't hide your business from ChatGPT search: OpenAI says "each setting is independent of the others"1 (more on GPTBot). Anthropic runs separate bots for Claude's search and its training2.
Google says Google-Extended "does not impact a site's inclusion in Google Search", where Googlebot is the control for AI Overviews and AI Mode4,5. It also controls "grounding" in the Gemini app and Vertex AI, answers built from pages in Google's index, so blocking it can keep you out of those.
Microsoft lists no separate training bot; under its 2023 rules, nocache limits training to title, link and snippet, and noarchive stops it but drops the page from Bing's AI answers13. Your call: a kuih maker with original recipes may opt out; we leave training open.
Do AI bots actually obey robots.txt?
AI bots that crawl automatically say they obey robots.txt; bots fetching a page for a user may not. OpenAI says robots.txt rules "may not apply" to ChatGPT-User, because a user starts the fetch1. Perplexity says Perplexity-User "generally ignores robots.txt rules"3, and Meta says its user fetcher "may bypass robots.txt rules"8.
In 2025, Cloudflare reported Perplexity using "a generic browser intended to impersonate Google Chrome" when its declared bots were blocked9. Perplexity replied that Cloudflare had misattributed a browser service's traffic10. Cloudflare saw ChatGPT-User obey a block in the same tests.
In 2026 TollBit saw scrapers "masquerading as other user agents, including Googlebot"14. Check a bot's address against the company's published IP list; as Cloudflare puts it, "robots.txt compliance is voluntary"11.
Which AI bots actually visit a small Malaysian website?
OpenAI's and Anthropic's crawlers visited easyfound.co most from 25 to 30 September 2026, as in our Bing check (per-bot data):
- ClaudeBot: 173 requests, mostly robots.txt and sitemap checks
- GPTBot: 117
- OAI-SearchBot: 49
- Googlebot: 47
- Bingbot: 2
All came from each company's published addresses; none came from Perplexity's, Apple's or Meta's AI bots.
Malaysian websites rarely block them: of 84 .my robots.txt files we read on 1 October, none fully blocked any AI bot in the table.
What should my robots.txt say to let AI search bots in?
Your robots.txt should allow search and answer bots, state your training choice and list your sitemap. Copy this into yourbusiness.com.my/robots.txt:
# Search and answer bots
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Googlebot
User-agent: Bingbot
User-agent: Applebot
User-agent: meta-webindexer
Allow: /
# Training (Meta's also indexes): Disallow to opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
Allow: /
User-agent: *
Allow: /
Sitemap: https://yourbusiness.com.my/sitemap.xml
- Named bots skip the
*group. Crawlers use it only "If no matching group exists"12. If your file blocks folders such as/wp-admin/, repeat those lines in each group. - Cloudflare can contradict or bypass this file. Its managed robots.txt adds its own
Disallow: /for GPTBot, ClaudeBot and other training bots11, and its AI bot block stops bots at the firewall (September 2026 defaults).
What can't our AI bot check tell me about my own website?
- We read files, not firewalls. A Cloudflare block wouldn't show in our robots.txt count.
- Our logs cover one small website. Busier sites may draw Perplexity, Apple or Meta.
- Bot names change. Copied lists still carry outdated ones such as
anthropic-ai.
How do I make sure AI tools can find my business once the right bots are in?
Letting the right bots in is free; then check the answers.
- Request EasyFound's free check. See what ChatGPT says about your business, then go through it with us on a free call.
- Compare your robots.txt with the table, or paste the template.
- On Cloudflare, check your AI bot policies and allow Search and Agent bots.
- Decide on training bots once.
EasyFound, an AI SEO service for Malaysian businesses, asks six AI tools about your business in English, Malay and Chinese, then works on the pages they read. Start with the free check below.
Which AI tools can find your business?
Letting the right bots in is step one. See what ChatGPT actually says about your business, then go through it with us on a free call.
Frequently asked questions
Should I block AI crawlers on my business website?
Don't block the bots that power AI search and answers, such as OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot and Applebot, or AI tools can't show your pages in their answers. Training crawlers such as GPTBot, ClaudeBot and Google-Extended are a separate choice that each business makes for itself. In EasyFound's check on 1 October 2026, none of 84 Malaysian .my websites' robots.txt files fully blocked any of these AI bots.
Should I allow ChatGPT-User and Perplexity-User if they ignore robots.txt anyway?
Yes, allowing them costs nothing and states your preference clearly. They open pages when someone asks ChatGPT or Perplexity a question that needs them, rather than crawling on their own. OpenAI says robots.txt rules 'may not apply' to ChatGPT-User, and Perplexity says Perplexity-User 'generally ignores robots.txt rules', so a Disallow line may not stop them; Anthropic says its Claude-User bot does follow robots.txt.
Does blocking Google-Extended remove my business from Google's AI Overviews?
No. Google says Google-Extended 'does not impact a site's inclusion in Google Search', and that robots.txt rules for Googlebot are the control for Search, which includes AI Overviews and AI Mode. Google-Extended does control two other things: training future Gemini models, and grounding answers in the Gemini app and Google's Vertex AI with content from Google's search index.
Do AI bots really follow robots.txt?
The automatic crawlers say they do: OpenAI, Anthropic, Google, Apple and Perplexity all document robots.txt controls for them. Fetchers that open a page because a user asked are different. OpenAI says robots.txt rules 'may not apply' to ChatGPT-User, and Perplexity says Perplexity-User 'generally ignores' them. In 2025 Cloudflare reported Perplexity crawling under undeclared identities, which Perplexity disputed. Cloudflare itself notes that robots.txt compliance is voluntary.
Sources
Our data: EasyFound, AI bots visiting easyfound.co, and AI bots blocked in Malaysian websites' robots.txt. Server logs: every request that reached easyfound.co from 25 September 2026, 15:35 UTC, to 30 September, 16:21 UTC. That is 3,360 requests in Railway's HTTP logs, deduplicated, and the same window as our Bing article. Our llms.txt article used a window ending 30 September, 02:48 UTC, so its counts for OAI-SearchBot (28), GPTBot (116) and Googlebot (28) are lower; it also counted the Claude Code fetch below as Claude-User. Bots were matched by user agent and checked against each company's published IP list, fetched 1 October 2026. Cloudflare-cached requests aren't logged, so counts are a floor. One request labelled Claude-User came from Claude Code, a developer tool that fetches from the user's own computer, so we left it out. robots.txt check: on 1 October 2026 we re-fetched the 167 sites from our 26 September check, which could read only 59 files; this re-check supersedes that count. This time we read 142 files. They include 84 on .my domains, and 36 from a list of 41 Malaysian business websites. The .my sites skew towards ones ChatGPT already cites. Download the results (CSV). We don't name the websites.
The ChatGPT answers mentioned come from asking this article's question three times through OpenAI's API (gpt-5.4-mini, web search available, location Malaysia), not the ChatGPT app, on 1 October 2026.
Other sources (all read 1 October 2026):
- OpenAI, Overview of OpenAI Crawlers (link)
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?, updated 7 April 2026 (link)
- Perplexity, Perplexity Crawlers (link)
- Google, List of Google's common crawlers, updated 14 July 2026 (link)
- Google Search Central, AI features and your website (link)
- Microsoft Bing, Which crawlers does Bing use? (link)
- Apple, About Applebot, 4 September 2026 (link)
- Meta, Meta Web Crawlers, updated 21 May 2026 (link)
- Cloudflare, Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives, 4 August 2025 (link)
- Perplexity, Agents or Bots? Making Sense of AI on the Open Web, 2025 (link)
- Cloudflare, robots.txt setting (managed robots.txt), updated 3 August 2026 (link)
- IETF, RFC 9309: Robots Exclusion Protocol, 2022 (link)
- Microsoft Bing (Fabrice Canel), Announcing new options for webmasters to control usage of their content in Bing Chat, 22 September 2023 (link)
- TollBit, State of the Bots, 2026 Q1 & Q2 (link)
The data is shared under CC BY 4.0: you're free to reuse it, with credit to EasyFound.
Bot names and company rules change; everything above was checked on 1 October 2026.
Research by Janice Pang, EasyFound.