ai-crawler-index 1.2.0 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +2 -2
- data/lib/ai_crawler_index/data.json +1 -1
- metadata +3 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 7eb405f35785eaba5bd089ca50e40baebcb30b9c9a51e9df0e4292bb276c6ebc
|
|
4
|
+
data.tar.gz: 406d4d37a363db59eef9225f320d0a7de0a2893dddc89eea8809714d90fec425
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: e1c32d3e957ad77f17a11f107264f18fdf8148dc9cae97b8897e45515ef9443aa2a0c7b517d7cb5f80396b28c52c3095e7a29460f1854466b1cc3b353f78ea70
|
|
7
|
+
data.tar.gz: e1777f672e750e268dd76114398f50d81d28b06973b8b581104b48195caa952145dee85b1763c8fd13b1394ac77e42dd3602cb810b923b9ef59eb2265d9023e0
|
data/README.md
CHANGED
|
@@ -98,14 +98,14 @@ carries an `ip_ranges` URL — check the address before you act on the name.
|
|
|
98
98
|
|
|
99
99
|
## The data
|
|
100
100
|
|
|
101
|
-
Generated 2026-09-
|
|
101
|
+
Generated 2026-09-05T17:25:28+00:00 from the [AI Crawler Index](https://www.pathwren.workers.dev/c/rubygems-registry/) —
|
|
102
102
|
150 crawlers from 74 operators, each reviewed against its
|
|
103
103
|
operator's own published documentation. Robots tokens, user-agent strings and
|
|
104
104
|
documentation URLs come from those operator pages (cited per record); the
|
|
105
105
|
categories and the prose are the index's own.
|
|
106
106
|
|
|
107
107
|
- Source of truth: `https://www.pathwren.workers.dev/c/rubygems-registry/data/agents.json` — regenerated every six hours.
|
|
108
|
-
- This bundle is a **snapshot of that file taken at 2026-09-
|
|
108
|
+
- This bundle is a **snapshot of that file taken at 2026-09-05T17:25:28+00:00**, not a
|
|
109
109
|
live feed. Crawlers appear and change names; a gem published last month cannot
|
|
110
110
|
know about a bot announced last week.
|
|
111
111
|
- Data licence: **CC0-1.0**. Code licence: MIT.
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":"1.2.0","generated_at":"2026-09-02T20:41:24+00:00","source":"AI Crawler Index — https://www.pathwren.workers.dev — independent, non-commercial; data CC0-1.0","data_url":"https://www.pathwren.workers.dev/c/rubygems-registry/data/agents.json","window":"Table reviewed 2026-09-02; every record checked against its operator's own published documentation. IP-range mirrors behind the index refresh every six hours; this bundle is a snapshot, not a live feed.","license":{"code":"MIT","data":"CC0-1.0"},"categories":{"ai-training":{"label":"AI training crawlers","description":"Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today."},"ai-search":{"label":"AI search crawlers","description":"Build the retrieval index an assistant answers and cites from. These are the crawlers that send you traffic; blocking them is the expensive mistake in this space."},"user-fetch":{"label":"User-triggered fetchers","description":"Fetch one page because a person asked for it, right then. One human intent, one request. Blocking them produces a visible error for a real reader."},"dataset":{"label":"Corpus and dataset builders","description":"Crawl the web into a published or resold dataset that other people train on. Highest leverage per block, longest delay before any effect."},"search":{"label":"Search engines","description":"Classic index-and-rank crawlers. Several also feed their operator's generative answers, which is why the AI opt-out for Google and Apple is a token rather than a block."},"seo":{"label":"SEO and backlink crawlers","description":"Commercial link-graph tooling. No user-facing effect either way, and usually a large share of your bot bandwidth."},"archive":{"label":"Archivers","description":"Preservation crawlers. Their output is public and permanent, which makes them a separate decision from the AI one."},"preview":{"label":"Link preview fetchers","description":"Read your Open Graph tags when someone shares a link. Blocking these is almost always an accident."},"tool":{"label":"Tools and frameworks","description":"Not operators: crawling software anyone can run. The party behind the request is unknown, so treat them as a rate-limit question rather than a consent question."}},"ai_categories":["ai-training","ai-search","user-fetch","dataset"],"regex":{"all":"(AdsBot\\-Google|AdsBot\\-Google\\-Mobile|AdsBot\\-Google\\-Mobile\\-Apps|AhrefsBot|AhrefsSiteAudit|AI2Bot|Ai2Bot\\-Dolma|aiHitBot|AIWebIndex|Amazonbot|Andibot|Anomura|anthropic\\-ai|APIs\\-Google|Applebot|archive\\.org_bot|atlassian\\-bot|AwarioRssBot|AwarioSmartBot|Baiduspider|barkrowler|bedrockbot|bingbot|Bytespider|CCBot|ChatGPT\\ Agent|ChatGPT\\-User|Claude\\-SearchBot|Claude\\-User|Claude\\-Web|ClaudeBot|Cloudflare\\-AutoRAG|cohere\\-ai|cohere\\-training\\-data\\-crawler|Cotoyogi|Crawl4AI|Crawlspace|DataForSeoBot|Diffbot|dotbot|DuckAssistBot|DuckDuckBot|EchoboxBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FeedFetcher\\-Google|FirecrawlAgent|Google\\-Agent|Google\\-CloudVertexBot|Google\\-CWS|Google\\-GeminiNotebook|Google\\-InspectionTool|Google\\-Pinpoint|Google\\-Read\\-Aloud|Google\\-Safety|Google\\-Site\\-Verification|Googlebot|Googlebot\\-Image|Googlebot\\-News|Googlebot\\-Video|GoogleMessages|GoogleOther|GoogleOther\\-Image|GoogleOther\\-Video|GoogleProducer|GPTBot|ia_archiver|ICC\\-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kagibot|KlaviyoAIBot|LAIONDownloader|Lightpanda|Linguee\\ Bot|Mediapartners\\-Google|meta\\-externalagent|meta\\-externalfetcher|Meta\\-WebIndexer|MistralAI\\-User|MJ12bot|MojeekBot|OAI\\-SearchBot|omgili|omgilibot|panscient\\.com|Perplexity\\-User|PerplexityBot|PetalBot|PhindBot|Pinterestbot|Poseidon\\ Research\\ Crawler|QualifiedBot|QuillBot|Qwantbot|Qwantbot\\-news|Reflectionbot|rogerbot|SBIntuitionsBot|Scrapy|Screaming\\ Frog\\ SEO\\ Spider|SemrushBot|SemrushBot\\-BA|SemrushBot\\-ESI|SemrushBot\\-FT|SemrushBot\\-OCOB|SemrushBot\\-SI|SemrushBot\\-SWA|SEOkicks|serpstatbot|SeznamBot|ShapBot|Sidetrade\\ indexer\\ bot|SiteAuditBot|Slackbot|Slackbot\\-LinkExpanding|SplitSignalBot|Storebot\\-Google|TerraCotta|Thinkbot|TikTokSpider|Timpibot|VelenPublicWebCrawler|Webzio\\-Extended|wpbot|YaK|YandexAdditional|YandexAdditionalBot|YandexBlogs|YandexBot|YandexCalendar|YandexComBot|YandexDirect|YandexFavicons|YandexImages|YandexMarket|YandexMedia|YandexMetrika|YandexMobileBot|YandexRenderResourcesBot|YandexScreenshotBot|YandexVideo|YandexWebmaster|Yeti|YouBot)","ai_only":"(AI2Bot|Ai2Bot\\-Dolma|aiHitBot|AIWebIndex|Amazonbot|Andibot|Anomura|anthropic\\-ai|atlassian\\-bot|AwarioRssBot|AwarioSmartBot|bedrockbot|Bytespider|CCBot|ChatGPT\\ Agent|ChatGPT\\-User|Claude\\-SearchBot|Claude\\-User|Claude\\-Web|ClaudeBot|Cloudflare\\-AutoRAG|cohere\\-ai|cohere\\-training\\-data\\-crawler|Cotoyogi|Diffbot|DuckAssistBot|EchoboxBot|ExaSearchBot|FacebookBot|Factset_spyderbot|Google\\-Agent|Google\\-CloudVertexBot|Google\\-GeminiNotebook|Google\\-Pinpoint|Google\\-Read\\-Aloud|GoogleOther|GoogleOther\\-Image|GoogleOther\\-Video|GPTBot|ICC\\-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|KlaviyoAIBot|LAIONDownloader|Linguee\\ Bot|meta\\-externalagent|meta\\-externalfetcher|Meta\\-WebIndexer|MistralAI\\-User|OAI\\-SearchBot|omgili|omgilibot|panscient\\.com|Perplexity\\-User|PerplexityBot|PhindBot|Poseidon\\ Research\\ Crawler|QualifiedBot|QuillBot|Reflectionbot|SBIntuitionsBot|SemrushBot\\-OCOB|ShapBot|Sidetrade\\ indexer\\ bot|TerraCotta|Thinkbot|TikTokSpider|VelenPublicWebCrawler|Webzio\\-Extended|YaK|YandexAdditional|YandexAdditionalBot|YandexCalendar|YouBot)","by_category":{"ai-training":"(anthropic\\-ai|Bytespider|ClaudeBot|cohere\\-training\\-data\\-crawler|Cotoyogi|FacebookBot|Factset_spyderbot|GoogleOther|GoogleOther\\-Image|GoogleOther\\-Video|GPTBot|ICC\\-Crawler|ISSCyberRiskCrawler|Linguee\\ Bot|meta\\-externalagent|Poseidon\\ Research\\ Crawler|QuillBot|Reflectionbot|SBIntuitionsBot|SemrushBot\\-OCOB|Sidetrade\\ indexer\\ bot|TikTokSpider|Webzio\\-Extended|YandexAdditional|YandexAdditionalBot)","ai-search":"(AIWebIndex|Amazonbot|Andibot|Anomura|atlassian\\-bot|bedrockbot|Claude\\-SearchBot|Claude\\-Web|Cloudflare\\-AutoRAG|DuckAssistBot|ExaSearchBot|Google\\-CloudVertexBot|KlaviyoAIBot|Meta\\-WebIndexer|OAI\\-SearchBot|PerplexityBot|PhindBot|QualifiedBot|ShapBot|TerraCotta|YouBot)","user-fetch":"(ChatGPT\\ Agent|ChatGPT\\-User|Claude\\-User|cohere\\-ai|Google\\-Agent|Google\\-GeminiNotebook|Google\\-Pinpoint|Google\\-Read\\-Aloud|meta\\-externalfetcher|MistralAI\\-User|Perplexity\\-User|YandexCalendar)","search":"(Applebot|Baiduspider|bingbot|DuckDuckBot|Googlebot|Googlebot\\-Image|Googlebot\\-News|Googlebot\\-Video|Kagibot|MojeekBot|PetalBot|Pinterestbot|Qwantbot|Qwantbot\\-news|SeznamBot|Storebot\\-Google|Timpibot|YandexBlogs|YandexBot|YandexComBot|YandexFavicons|YandexImages|YandexMarket|YandexMedia|YandexMobileBot|YandexRenderResourcesBot|YandexVideo|Yeti)","tool":"(AdsBot\\-Google|AdsBot\\-Google\\-Mobile|AdsBot\\-Google\\-Mobile\\-Apps|APIs\\-Google|Crawl4AI|Crawlspace|FeedFetcher\\-Google|FirecrawlAgent|Google\\-CWS|Google\\-InspectionTool|Google\\-Safety|Google\\-Site\\-Verification|GoogleProducer|Lightpanda|Mediapartners\\-Google|Scrapy|Screaming\\ Frog\\ SEO\\ Spider|wpbot|YandexDirect|YandexMetrika|YandexScreenshotBot|YandexWebmaster)","dataset":"(AI2Bot|Ai2Bot\\-Dolma|aiHitBot|AwarioRssBot|AwarioSmartBot|CCBot|Diffbot|EchoboxBot|ImagesiftBot|img2dataset|LAIONDownloader|omgili|omgilibot|panscient\\.com|Thinkbot|VelenPublicWebCrawler|YaK)","preview":"(facebookexternalhit|GoogleMessages|Slackbot|Slackbot\\-LinkExpanding)","seo":"(AhrefsBot|AhrefsSiteAudit|barkrowler|DataForSeoBot|dotbot|MJ12bot|rogerbot|SemrushBot|SemrushBot\\-BA|SemrushBot\\-ESI|SemrushBot\\-FT|SemrushBot\\-SI|SemrushBot\\-SWA|SEOkicks|serpstatbot|SiteAuditBot|SplitSignalBot)","archive":"(archive\\.org_bot|ia_archiver)"}},"crawlers":[{"slug":"adsbot-google","name":"AdsBot-Google","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"AdsBot-Google","robots_token":"AdsBot-Google","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"own-token-only","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/adsbot-google.html","what_it_is":"Checks the quality of desktop landing pages for Google Ads. Google documents that it ignores the robots.txt * group with the ad publisher's permission, and obeys a group named for its own token.","cost_of_blocking":"Google Ads cannot score your landing pages, which lowers Ad Rank on the ads pointing at them. If you do not buy ads, blocking it costs nothing but bandwidth savings."},{"slug":"adsbot-google-mobile","name":"AdsBot-Google-Mobile","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"AdsBot-Google-Mobile","robots_token":"AdsBot-Google-Mobile","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"own-token-only","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/adsbot-google-mobile.html","what_it_is":"The mobile-web landing page checker for Google Ads. Same rules as AdsBot-Google: the * group does not apply to it, its own token does.","cost_of_blocking":"Mobile ad landing pages go unscored and the ads pointing at them rank worse. No effect on organic search."},{"slug":"adsbot-google-mobile-apps","name":"AdsBot-Google-Mobile-Apps","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"AdsBot-Google-Mobile-Apps","robots_token":"AdsBot-Google-Mobile-Apps","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"own-token-only","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/adsbot-google-mobile-apps.html","what_it_is":"Checks Android app landing pages for Google Ads. It obeys a group named for its own token and, per Google, follows the AdsBot-Google rules otherwise.","cost_of_blocking":"App-install ad landing pages go unscored. Nothing organic changes."},{"slug":"ahrefsbot","name":"AhrefsBot","operator":"Ahrefs","category":"seo","category_label":"SEO and backlink crawlers","ua":"AhrefsBot","robots_token":"AhrefsBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://ahrefs.com/robot","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/ahrefs-crawler.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ahrefsbot.html","what_it_is":"Ahrefs' backlink crawler, and one of the largest non-search crawlers on the web by request volume.","cost_of_blocking":"No user-facing effect. Ahrefs honours Crawl-delay, so rate-limiting is usually better than blocking."},{"slug":"ahrefssiteaudit","name":"AhrefsSiteAudit","operator":"Ahrefs","category":"seo","category_label":"SEO and backlink crawlers","ua":"AhrefsSiteAudit","robots_token":"AhrefsSiteAudit","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://ahrefs.com/robot","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/ahrefs-crawler.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ahrefssiteaudit.html","what_it_is":"Ahrefs' site-audit crawler, separate from AhrefsBot. Ahrefs documents that it obeys robots.txt by default, and that a verified site owner can ask for it to be allowed to ignore robots.txt on their own site so the audit can see disallowed sections.","cost_of_blocking":"Site owners auditing your domain get an incomplete report. Blocking it saves bandwidth and costs you nothing in search."},{"slug":"ai2bot","name":"AI2Bot","operator":"Allen Institute for AI","category":"dataset","category_label":"Corpus and dataset builders","ua":"AI2Bot","robots_token":"AI2Bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://allenai.org/crawler","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ai2bot.html","what_it_is":"The Allen Institute's crawler, gathering pages for open research corpora such as Dolma that underpin fully open models like OLMo.","cost_of_blocking":"Excluded from open research datasets. Worth a deliberate decision: this is the category where 'blocking AI' also blocks the open, auditable end of it."},{"slug":"ai2bot-dolma","name":"Ai2Bot-Dolma","operator":"Allen Institute for AI","category":"dataset","category_label":"Corpus and dataset builders","ua":"Ai2Bot-Dolma","robots_token":"Ai2Bot-Dolma","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://allenai.org/crawler","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ai2bot-dolma.html","what_it_is":"The variant of AI2's crawler named for the Dolma corpus specifically.","cost_of_blocking":"Same as AI2Bot: exclusion from an open, published training corpus."},{"slug":"aihitbot","name":"aiHitBot","operator":"aiHit","category":"dataset","category_label":"Corpus and dataset builders","ua":"aiHitBot","robots_token":"aiHitBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.aihitdata.com/about","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/aihitbot.html","what_it_is":"aiHit's automated collector, building a company dataset from public company websites.","cost_of_blocking":"Your company record in a B2B dataset goes stale. Documented as respecting robots.txt, so the rule works."},{"slug":"aiwebindex","name":"AIWebIndex","operator":"Lyrenth","category":"ai-search","category_label":"AI search crawlers","ua":"AIWebIndex","robots_token":"AIWebIndex","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://lyrenth.com/crawler-policy","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/aiwebindex.html","what_it_is":"Builds an index of public pages and serves them to AI agents as extracted readable text, with attribution and a link back. Lyrenth publishes a crawler policy stating it does not train foundation models on what it collects and that it obeys robots.txt.","cost_of_blocking":"Agents reading through this index stop seeing you — including the attribution and link back that make it a referral rather than a summary."},{"slug":"amazonbot","name":"Amazonbot","operator":"Amazon","category":"ai-search","category_label":"AI search crawlers","ua":"Amazonbot","robots_token":"Amazonbot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://developer.amazon.com/amazonbot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/amazonbot.html","what_it_is":"Amazon's crawler, feeding Alexa's ability to answer questions from the web and Amazon's own search and assistant products.","cost_of_blocking":"Alexa and Amazon's assistants stop answering from your pages. Verify with reverse DNS to crawl.amazonbot.amazon before trusting the user-agent."},{"slug":"andibot","name":"Andibot","operator":"Andi","category":"ai-search","category_label":"AI search crawlers","ua":"Andibot","robots_token":"Andibot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://andisearch.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/andibot.html","what_it_is":"The crawler for Andi, a small generative search assistant that summarises pages rather than listing them.","cost_of_blocking":"You disappear from another assistant's answers. Andi publishes no robots.txt statement."},{"slug":"anomura","name":"Anomura","operator":"Direqt","category":"ai-search","category_label":"AI search crawlers","ua":"Anomura","robots_token":"Anomura","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://direqt.ai","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/anomura.html","what_it_is":"Direqt's search crawler. It indexes the sites of Direqt's own publisher customers so their on-site chatbots can answer from them.","cost_of_blocking":"If you are the publisher, this breaks the assistant you put on your own pages. If you are not, it should not be crawling you."},{"slug":"anthropic-ai","name":"anthropic-ai","operator":"Anthropic","category":"ai-training","category_label":"AI training crawlers","ua":"anthropic-ai","robots_token":"anthropic-ai","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"n-a","docs":"https://support.anthropic.com/en/articles/8896518","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/anthropic-ai.html","what_it_is":"A legacy robots.txt token from before Anthropic consolidated on ClaudeBot. It is still widely present in robots.txt files and costs nothing to keep, but it is a control token rather than a bot you will see in logs.","cost_of_blocking":"None. Nothing crawls under this name today; keeping the rule is harmless insurance."},{"slug":"apis-google","name":"APIs-Google","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"APIs-Google","robots_token":"APIs-Google","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"own-token-only","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/apis-google.html","what_it_is":"Delivers push notifications for Google APIs to a webhook you registered. It is a special-case crawler: it ignores the robots.txt * group, because the fetch is a delivery to an address you asked it to deliver to.","cost_of_blocking":"Google API push notifications stop arriving at your endpoint. This only affects services you set up yourself; there is no search or AI consequence."},{"slug":"applebot","name":"Applebot","operator":"Apple","category":"search","category_label":"Search engines","ua":"Applebot","robots_token":"Applebot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://support.apple.com/en-us/119829","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/apple-applebot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/applebot.html","what_it_is":"Powers Siri, Spotlight and Safari suggestions. Blocking it is a search decision, not an AI decision — the AI decision has its own token.","cost_of_blocking":"You disappear from Siri, Spotlight and Safari search suggestions across Apple's install base."},{"slug":"applebot-extended","name":"Applebot-Extended","operator":"Apple","category":"ai-training","category_label":"AI training crawlers","ua":"(control token only — no crawler)","robots_token":"Applebot-Extended","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"n-a","docs":"https://support.apple.com/en-us/119829","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/applebot-extended.html","what_it_is":"Apple's counterpart to Google-Extended: a robots.txt token that withdraws consent for Apple Intelligence and Apple foundation-model training, without touching Applebot's search crawl.","cost_of_blocking":"Excluded from Apple Intelligence training. Siri, Spotlight and Safari suggestions are unaffected."},{"slug":"archive-org-bot","name":"archive.org_bot","operator":"Internet Archive","category":"archive","category_label":"Archivers","ua":"archive.org_bot","robots_token":"archive.org_bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://archive.org/details/archive.org_bot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/archive-org-bot.html","what_it_is":"The Wayback Machine's crawler. Preservation rather than AI, but it lands in the same 'is this bot welcome' decision and its output is a public corpus.","cost_of_blocking":"Your site stops being preserved. When it dies, it is gone. Consider this one separately from the AI question."},{"slug":"atlassian-bot","name":"atlassian-bot","operator":"Atlassian","category":"ai-search","category_label":"AI search crawlers","ua":"atlassian-bot","robots_token":"atlassian-bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://support.atlassian.com/organization-administration/docs/connect-custom-website-to-rovo/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/atlassian-bot.html","what_it_is":"Indexes a website so it can be searched and cited by Rovo, Atlassian's generative assistant inside Jira and Confluence. Atlassian's documentation walks a customer through editing robots.txt for it, which is as close to a compliance statement as this list gets.","cost_of_blocking":"Rovo cannot answer from your public documentation. If your customers live inside Atlassian tools, this is a support-deflection block."},{"slug":"awariorssbot","name":"AwarioRssBot","operator":"Awario","category":"dataset","category_label":"Corpus and dataset builders","ua":"AwarioRssBot","robots_token":"AwarioRssBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://awario.com/bots.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/awariorssbot.html","what_it_is":"The feed-reading half of Awario's pair, documented on the same page and under the same crawl-rate policy.","cost_of_blocking":"Your RSS updates stop reaching Awario's monitoring. Block both tokens or neither."},{"slug":"awariosmartbot","name":"AwarioSmartBot","operator":"Awario","category":"dataset","category_label":"Corpus and dataset builders","ua":"AwarioSmartBot","robots_token":"AwarioSmartBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://awario.com/bots.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/awariosmartbot.html","what_it_is":"Awario's brand-monitoring crawler. It documents one request per three seconds, honours Crawl-delay, and states it does not use consecutive IP blocks so identification is by user-agent only.","cost_of_blocking":"Mentions of brands on your pages stop being surfaced to the people monitoring them — including, quite possibly, your own."},{"slug":"baiduspider","name":"Baiduspider","operator":"Baidu","category":"search","category_label":"Search engines","ua":"Baiduspider","robots_token":"Baiduspider","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://help.baidu.com/question?prod_id=99&class=0&id=3001","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/baiduspider.html","what_it_is":"Baidu's search crawler, and the ingest path for Baidu's Ernie-backed answers.","cost_of_blocking":"Removal from Baidu Search, which matters only if you want Chinese-language traffic."},{"slug":"barkrowler","name":"Barkrowler","operator":"Babbar","category":"seo","category_label":"SEO and backlink crawlers","ua":"barkrowler","robots_token":"barkrowler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://babbar.tech/crawler","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/barkrowler.html","what_it_is":"Babbar's crawler, which builds the link graph behind their French-market SEO tooling.","cost_of_blocking":"You leave Babbar's index. No effect on search or assistants."},{"slug":"bedrockbot","name":"bedrockbot","operator":"Amazon","category":"ai-search","category_label":"AI search crawlers","ua":"bedrockbot","robots_token":"bedrockbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/webcrawl-data-source-connector.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/bedrockbot.html","what_it_is":"The web crawler an AWS customer points at URLs they chose, to build a knowledge base for a Bedrock application. AWS documents that it respects robots.txt and that the user-agent carries a per-customer suffix, so you can allow or refuse one customer's crawl by naming bedrockbot-UUID.","cost_of_blocking":"Companies building retrieval applications on Bedrock cannot include your pages. This is a RAG block, not a training block: nothing is being trained, but nothing can cite you either."},{"slug":"bingbot","name":"bingbot","operator":"Microsoft","category":"search","category_label":"Search engines","ua":"bingbot","robots_token":"bingbot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/bing-bingbot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/bingbot.html","what_it_is":"Bing's only crawler, and therefore also the crawler behind Microsoft Copilot's grounding. Microsoft's documented way to keep search indexing while refusing generative reuse is the nocache / noarchive robots meta directive, not a separate user-agent.","cost_of_blocking":"Very high and very wide: Bing, Copilot, DuckDuckGo and several assistants that resell Bing's index all lose you at once. Use nocache/noarchive rather than blocking."},{"slug":"bytespider","name":"Bytespider","operator":"ByteDance","category":"ai-training","category_label":"AI training crawlers","ua":"Bytespider","robots_token":"Bytespider","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"disputed","docs":"https://www.bytespider.net/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/bytespider.html","what_it_is":"ByteDance's crawler, associated with training data collection for Doubao and related models. Repeatedly reported by CDNs and site operators as the highest-volume AI crawler on the web and as inconsistent about robots.txt.","cost_of_blocking":"Little to lose. If you want it gone, expect to block by user-agent at the edge rather than to ask politely in robots.txt."},{"slug":"ccbot","name":"CCBot","operator":"Common Crawl","category":"dataset","category_label":"Corpus and dataset builders","ua":"CCBot","robots_token":"CCBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://commoncrawl.org/ccbot","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/commoncrawl-ccbot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ccbot.html","what_it_is":"Common Crawl's corpus builder. It trains nothing itself, but its archive is an input to most open and many closed LLM training sets, which makes it the highest-leverage single entry on this list.","cost_of_blocking":"Future Common Crawl snapshots exclude you, so downstream training sets lose you too — but only going forward. Existing snapshots are permanent and blocking today does not retract them."},{"slug":"chatgpt-agent","name":"ChatGPT Agent","operator":"OpenAI","category":"user-fetch","category_label":"User-triggered fetchers","ua":"ChatGPT Agent","robots_token":"ChatGPT-User","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://platform.openai.com/docs/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/openai-chatgpt-user.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/chatgpt-agent.html","what_it_is":"ChatGPT's agent mode driving a real browser: it navigates and interacts with sites to finish a multi-step task a user gave it. OpenAI governs it with the ChatGPT-User token and the ChatGPT-User prefix list rather than a token of its own, so the robots rule and the address check are the same ones.","cost_of_blocking":"Agentic tasks a user asked for — booking, comparing, filling a form on your site — fail. This is the fetch that ends in a transaction, so it is the most expensive user-triggered block on this list."},{"slug":"chatgpt-user","name":"ChatGPT-User","operator":"OpenAI","category":"user-fetch","category_label":"User-triggered fetchers","ua":"ChatGPT-User","robots_token":"ChatGPT-User","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://platform.openai.com/docs/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/openai-chatgpt-user.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/chatgpt-user.html","what_it_is":"Fetches a single page at the moment a user or a ChatGPT agent asks for it — a pasted link, a browsing step, an Operator task. One human intent, one request. OpenAI states these fetches are not used for training.","cost_of_blocking":"ChatGPT cannot open your pages when a user explicitly asks it to. The user sees a fetch failure. This is usually the last bot anyone means to block."},{"slug":"claude-searchbot","name":"Claude-SearchBot","operator":"Anthropic","category":"ai-search","category_label":"AI search crawlers","ua":"Claude-SearchBot","robots_token":"Claude-SearchBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://support.anthropic.com/en/articles/8896518","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/claude-searchbot.html","what_it_is":"Indexes pages so Claude's web search can find and cite them. Separate token from the training crawler, so search visibility and training consent are independent decisions.","cost_of_blocking":"You stop appearing in Claude's search results and citations."},{"slug":"claude-user","name":"Claude-User","operator":"Anthropic","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Claude-User","robots_token":"Claude-User","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://support.anthropic.com/en/articles/8896518","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/claude-user.html","what_it_is":"Fetches a page because a Claude user asked Claude to read it, at that moment.","cost_of_blocking":"Claude reports a fetch failure to a user who asked for your page by name."},{"slug":"claude-web","name":"Claude-Web","operator":"Anthropic","category":"ai-search","category_label":"AI search crawlers","ua":"Claude-Web","robots_token":"Claude-Web","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"n-a","docs":"https://support.anthropic.com/en/articles/8896518","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/claude-web.html","what_it_is":"An earlier Anthropic token for user-facing web access, superseded by Claude-User and Claude-SearchBot. Kept here because it appears in most published robots.txt templates.","cost_of_blocking":"None in practice. Retain the rule; expect no traffic."},{"slug":"claudebot","name":"ClaudeBot","operator":"Anthropic","category":"ai-training","category_label":"AI training crawlers","ua":"ClaudeBot","robots_token":"ClaudeBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://support.anthropic.com/en/articles/8896518","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/claudebot.html","what_it_is":"Anthropic's bulk crawler, gathering pages that may be used to train Claude models.","cost_of_blocking":"Content excluded from training data for future Claude models. No effect on Claude's ability to fetch a link a user gives it."},{"slug":"cloudflare-autorag","name":"Cloudflare-AutoRAG","operator":"Cloudflare","category":"ai-search","category_label":"AI search crawlers","ua":"Cloudflare-AutoRAG","robots_token":"Cloudflare-AutoRAG","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://developers.cloudflare.com/ai-search/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/cloudflare-autorag.html","what_it_is":"The crawler behind Cloudflare's AI Search / AutoRAG, which indexes a website into a retrieval index for an application. Cloudflare's own documentation warns that a bot-blocking rule on your zone will also stop this crawler and tells you to allow-list it.","cost_of_blocking":"Applications built on Cloudflare AI Search cannot retrieve your pages. If you are the one building the index over your own site, blocking it breaks your own product."},{"slug":"cohere-ai","name":"cohere-ai","operator":"Cohere","category":"user-fetch","category_label":"User-triggered fetchers","ua":"cohere-ai","robots_token":"cohere-ai","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://cohere.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/cohere-ai.html","what_it_is":"Cohere's fetcher, used when its assistant products need a page.","cost_of_blocking":"Cohere-powered assistants cannot read your pages on request."},{"slug":"cohere-training-data-crawler","name":"cohere-training-data-crawler","operator":"Cohere","category":"ai-training","category_label":"AI training crawlers","ua":"cohere-training-data-crawler","robots_token":"cohere-training-data-crawler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://cohere.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/cohere-training-data-crawler.html","what_it_is":"Cohere's separately-named bulk crawler for model training data, split out so consent for training and consent for retrieval can differ.","cost_of_blocking":"Excluded from Cohere model training."},{"slug":"cotoyogi","name":"Cotoyogi","operator":"ROIS-DS","category":"ai-training","category_label":"AI training crawlers","ua":"Cotoyogi","robots_token":"Cotoyogi","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://ds.rois.ac.jp/en_center8/en_crawler/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/cotoyogi.html","what_it_is":"A crawler run by ROIS-DS, a Japanese inter-university research organisation, collecting Japanese-language text for AI training. It publishes a crawler page in English and Japanese.","cost_of_blocking":"Your Japanese-language content is left out of an academic training corpus."},{"slug":"crawl4ai","name":"Crawl4AI","operator":"Crawl4AI project","category":"tool","category_label":"Tools and frameworks","ua":"Crawl4AI","robots_token":"Crawl4AI","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://github.com/unclecode/crawl4ai","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/crawl4ai.html","what_it_is":"An open-source LLM-oriented crawler and scraper library, run by whoever installs it. Like Scrapy, the default user-agent identifies the software and says nothing about who is behind the request.","cost_of_blocking":"You block a library, not an operator: the rule catches a researcher and a bulk scraper equally, and anyone who edits one config line is not caught at all."},{"slug":"crawlspace","name":"Crawlspace","operator":"Crawlspace","category":"tool","category_label":"Tools and frameworks","ua":"Crawlspace","robots_token":"Crawlspace","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://crawlspace.dev","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/crawlspace.html","what_it_is":"A crawling platform: customers run their own crawls on it to feed agents, RAG pipelines and structured-data workflows. Like Firecrawl, the party behind any given request is the customer, not the platform.","cost_of_blocking":"Whatever any Crawlspace customer was building over your pages stops working. Volume and intent vary per customer, so this is a rate-limit decision more than a consent one."},{"slug":"dataforseobot","name":"DataForSeoBot","operator":"DataForSEO","category":"seo","category_label":"SEO and backlink crawlers","ua":"DataForSeoBot","robots_token":"DataForSeoBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://dataforseo.com/dataforseo-bot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/dataforseobot.html","what_it_is":"Builds the backlink and SERP datasets DataForSEO resells through its API, so one crawl reaches many downstream tools.","cost_of_blocking":"You leave a dataset that a long tail of SEO products is built on. No user-facing effect."},{"slug":"diffbot","name":"Diffbot","operator":"Diffbot","category":"dataset","category_label":"Corpus and dataset builders","ua":"Diffbot","robots_token":"Diffbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://docs.diffbot.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/diffbot.html","what_it_is":"Extracts structured records from pages to build a commercial knowledge graph that is resold and used for retrieval and training.","cost_of_blocking":"Your facts stop entering a widely-licensed knowledge graph. Whether that is a loss depends on whether you want to be a machine-readable entity."},{"slug":"dotbot","name":"DotBot","operator":"Moz","category":"seo","category_label":"SEO and backlink crawlers","ua":"dotbot","robots_token":"dotbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://moz.com/help/moz-procedures/crawlers/dotbot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/dotbot.html","what_it_is":"Moz's crawler for Link Explorer. Moz documents that it respects robots.txt and that dotbot is the token to name.","cost_of_blocking":"You leave Moz's link index, so Domain Authority and link reports about your site get thinner. Nothing a reader or an assistant sees changes."},{"slug":"duckassistbot","name":"DuckAssistBot","operator":"DuckDuckGo","category":"ai-search","category_label":"AI search crawlers","ua":"DuckAssistBot","robots_token":"DuckAssistBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/duckassistbot.html","what_it_is":"Fetches pages so DuckAssist can generate and cite answers inside DuckDuckGo.","cost_of_blocking":"No DuckAssist answers or citations from your site. Ordinary DuckDuckGo results are unaffected."},{"slug":"duckduckbot","name":"DuckDuckBot","operator":"DuckDuckGo","category":"search","category_label":"Search engines","ua":"DuckDuckBot","robots_token":"DuckDuckBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://duckduckgo.com/duckduckgo-help-pages/results/duckduckbot/","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/duckduckgo-duckduckbot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/duckduckbot.html","what_it_is":"DuckDuckGo's own crawler. Note that the bulk of DuckDuckGo's web results come from Bing, so blocking bingbot removes you from DuckDuckGo whether or not you allow this one.","cost_of_blocking":"Limited on its own; the real DuckDuckGo lever is bingbot."},{"slug":"echoboxbot","name":"EchoboxBot","operator":"Echobox","category":"dataset","category_label":"Corpus and dataset builders","ua":"EchoboxBot","robots_token":"EchoboxBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://echobox.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/echoboxbot.html","what_it_is":"Collects data supporting Echobox's AI-driven social and email distribution products, which publishers use to schedule and target their own content.","cost_of_blocking":"Publishers using Echobox get worse scheduling decisions about your articles. No compliance statement is published."},{"slug":"exasearchbot","name":"ExaSearchBot","operator":"Exa","category":"ai-search","category_label":"AI search crawlers","ua":"ExaSearchBot","robots_token":"ExaSearchBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://exa.ai","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/exasearchbot.html","what_it_is":"Exa's crawler. It discovers and indexes public pages so they can be retrieved and cited through Exa's search API, which is one of the common retrieval backends behind agent frameworks.","cost_of_blocking":"Agents built on Exa's API stop finding you. Exa publishes no statement about robots.txt compliance, so treat the rule as a request."},{"slug":"facebookbot","name":"FacebookBot","operator":"Meta","category":"ai-training","category_label":"AI training crawlers","ua":"FacebookBot","robots_token":"FacebookBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/facebookbot.html","what_it_is":"Meta's older speech- and language-corpus crawler, largely superseded by meta-externalagent but still listed as a valid robots token.","cost_of_blocking":"Negligible today. Keep the rule; expect little traffic."},{"slug":"facebookexternalhit","name":"facebookexternalhit","operator":"Meta","category":"preview","category_label":"Link preview fetchers","ua":"facebookexternalhit","robots_token":"facebookexternalhit","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/facebookexternalhit.html","what_it_is":"The link unfurler: it reads your Open Graph tags when somebody shares your URL on a Meta property.","cost_of_blocking":"Severe and usually accidental. Your links share as bare grey boxes with no title, image or description across Facebook, Instagram, Messenger and WhatsApp. Almost nobody means to block this."},{"slug":"factset-spyderbot","name":"Factset_spyderbot","operator":"FactSet","category":"ai-training","category_label":"AI training crawlers","ua":"Factset_spyderbot","robots_token":"Factset_spyderbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.factset.com/ai","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/factset-spyderbot.html","what_it_is":"FactSet's crawler, collecting data used in AI model training for its financial data and analytics products.","cost_of_blocking":"Exclusion from a financial-data vendor's corpus. Relevant mostly to companies whose filings and disclosures are being read."},{"slug":"feedfetcher-google","name":"FeedFetcher-Google","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"FeedFetcher-Google","robots_token":"FeedFetcher-Google","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/feedfetcher-google.html","what_it_is":"Crawls RSS and Atom feeds for Google News and WebSub. It is a user-triggered fetcher, and Google documents that those generally ignore robots.txt because a person asked for the fetch. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"Feed-driven Google products stop seeing your updates. A robots.txt rule will not stop it — block by user-agent at the edge if you mean it."},{"slug":"firecrawlagent","name":"FirecrawlAgent","operator":"Firecrawl","category":"tool","category_label":"Tools and frameworks","ua":"FirecrawlAgent","robots_token":"FirecrawlAgent","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://docs.firecrawl.dev/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/firecrawlagent.html","what_it_is":"A hosted scrape-to-markdown service that LLM applications call to read pages. The requester is whoever is building on it, not Firecrawl itself, so volume and intent vary wildly.","cost_of_blocking":"Applications built on Firecrawl cannot read your pages. This is increasingly how agents fetch the web, so it is a bigger block than its name suggests."},{"slug":"google-agent","name":"Google-Agent","operator":"Google","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Google-Agent","robots_token":"Google-Agent","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered-agents.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-agent.html","what_it_is":"Agents hosted on Google infrastructure navigating the web and taking actions on a user's request. Google names one prefix list for it — user-triggered-agents.json — and is separately experimenting with Web Bot Auth under the identity https://agent.bot.goog.","cost_of_blocking":"Google-hosted agents cannot complete a task on your site for a user who asked them to. This is the agentic-commerce fetch: blocking it removes you from what an assistant can actually do rather than from what it can say."},{"slug":"google-cloudvertexbot","name":"Google-CloudVertexBot","operator":"Google","category":"ai-search","category_label":"AI search crawlers","ua":"Google-CloudVertexBot","robots_token":"Google-CloudVertexBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-cloudvertexbot.html","what_it_is":"Crawls a site on behalf of a Vertex AI Agent Builder customer who is building an agent over that site. It only visits sites the customer has asked it to.","cost_of_blocking":"Third parties can no longer build Vertex AI agents that read your site. Irrelevant to Google Search."},{"slug":"google-cws","name":"Google-CWS","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"Google-CWS","robots_token":"Google-CWS","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-cws.html","what_it_is":"The Chrome Web Store fetcher. It requests the URLs a developer put in the metadata of a Chrome extension or theme. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"Chrome Web Store listings that point at your pages cannot fetch them. Relevant only if you publish extensions."},{"slug":"google-extended","name":"Google-Extended","operator":"Google","category":"ai-training","category_label":"AI training crawlers","ua":"(control token only — no crawler)","robots_token":"Google-Extended","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"n-a","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-extended.html","what_it_is":"Not a crawler. A robots.txt token that tells Google whether pages Googlebot already fetched may be used to train and ground Gemini. You will never see it in an access log; disallowing it changes what Google does with content it fetched under a different name.","cost_of_blocking":"You are excluded from Gemini grounding and Gemini training. Google Search ranking and indexing are explicitly unaffected. This is the cleanest 'no training, keep my search traffic' lever that exists."},{"slug":"google-gemininotebook","name":"Google-GeminiNotebook","operator":"Google","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Google-GeminiNotebook","robots_token":"Google-GeminiNotebook","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-gemininotebook.html","what_it_is":"Fetches a URL a Gemini Notebook (formerly NotebookLM) user added as a source to their notebook. The former agent string Google-NotebookLM is documented as supported until August 2026. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"A user who deliberately added your page as a research source gets nothing. This is a citation-shaped fetch, not a training crawl."},{"slug":"google-inspectiontool","name":"Google-InspectionTool","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"Google-InspectionTool","robots_token":"Google-InspectionTool","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-inspectiontool.html","what_it_is":"The fetcher behind Search Console's URL Inspection and the Rich Results Test. It runs when a site owner clicks a button.","cost_of_blocking":"Your own Search Console live tests stop working. Blocking this only hurts you."},{"slug":"google-pinpoint","name":"Google-Pinpoint","operator":"Google","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Google-Pinpoint","robots_token":"Google-Pinpoint","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-pinpoint.html","what_it_is":"Fetches individual URLs that a Pinpoint user — usually a journalist or researcher — added as a source to their own document collection. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"A researcher who explicitly added your page to a collection cannot load it."},{"slug":"google-read-aloud","name":"Google-Read-Aloud","operator":"Google","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Google-Read-Aloud","robots_token":"Google-Read-Aloud","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-read-aloud.html","what_it_is":"Fetches a page so Google can read it out loud with text-to-speech, at the moment a user asks. Formerly google-speakr. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"A reader who asked Google to read your page aloud — often someone using it for accessibility — gets an error instead."},{"slug":"google-safety","name":"Google-Safety","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"Google-Safety","robots_token":"Google-Safety","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-safety.html","what_it_is":"Google's abuse-investigation fetcher: malware review, phishing reports and similar. Google documents that it ignores robots.txt entirely, and a robots.txt rule for it does nothing.","cost_of_blocking":"Nothing you can control. The rule is ignored by design; listing the token is documentation, not enforcement."},{"slug":"google-site-verification","name":"Google-Site-Verification","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"Google-Site-Verification","robots_token":"Google-Site-Verification","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-site-verification.html","what_it_is":"Fetches the token file or meta tag that proves you own a site, when you click verify in Search Console. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"Your own Search Console verification fails. Blocking this only ever hurts the person doing the blocking."},{"slug":"googlebot","name":"Googlebot","operator":"Google","category":"search","category_label":"Search engines","ua":"Googlebot","robots_token":"Googlebot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-googlebot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googlebot.html","what_it_is":"The classic search crawler. It is also the crawler behind AI Overviews: Google does not run a separate bot for them, which is why the only AI opt-out is the Google-Extended token and not a Googlebot block.","cost_of_blocking":"Total. You leave Google Search. Never block this to avoid AI use; use Google-Extended instead."},{"slug":"googlebot-image","name":"Googlebot-Image","operator":"Google","category":"search","category_label":"Search engines","ua":"Googlebot-Image","robots_token":"Googlebot-Image","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-googlebot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googlebot-image.html","what_it_is":"Image indexing for Google Images. A separate token so you can leave images out of search without leaving search.","cost_of_blocking":"Your images stop appearing in Google Images."},{"slug":"googlebot-news","name":"Googlebot-News","operator":"Google","category":"search","category_label":"Search engines","ua":"Googlebot-News","robots_token":"Googlebot-News","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-googlebot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googlebot-news.html","what_it_is":"A robots.txt token controlling inclusion in Google News. It does not have its own user-agent string; the fetch arrives as Googlebot.","cost_of_blocking":"Removal from Google News, with normal Search unaffected."},{"slug":"googlebot-video","name":"Googlebot-Video","operator":"Google","category":"search","category_label":"Search engines","ua":"Googlebot-Video","robots_token":"Googlebot-Video","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-googlebot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googlebot-video.html","what_it_is":"The video half of Googlebot. It crawls video files and the pages around them for Google Video search, and it is matched by a robots.txt group for Googlebot as well as by its own token.","cost_of_blocking":"Your videos leave Google video search. A rule for Googlebot already covers it, so blocking this token alone is usually a mistake of precision rather than of intent."},{"slug":"googlemessages","name":"GoogleMessages","operator":"Google","category":"preview","category_label":"Link preview fetchers","ua":"GoogleMessages","robots_token":"GoogleMessages","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googlemessages.html","what_it_is":"Generates the link preview when somebody sends one of your URLs in Google Messages. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"Your links appear as bare URLs with no title or image in Google Messages chats. A preview fetcher is almost never the one you meant to block."},{"slug":"googleother","name":"GoogleOther","operator":"Google","category":"ai-training","category_label":"AI training crawlers","ua":"GoogleOther","robots_token":"GoogleOther","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googleother.html","what_it_is":"A generic fetcher used by Google product teams for one-off crawls and research, including data collection that does not belong to Search.","cost_of_blocking":"No effect on Search indexing. Blocks internal Google research and product fetches."},{"slug":"googleother-image","name":"GoogleOther-Image","operator":"Google","category":"ai-training","category_label":"AI training crawlers","ua":"GoogleOther-Image","robots_token":"GoogleOther-Image","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googleother-image.html","what_it_is":"The image variant of GoogleOther: one-off fetches by Google product and research teams that are not Search. It also answers to a GoogleOther group in robots.txt.","cost_of_blocking":"Google teams outside Search stop fetching your images. Image Search itself is unaffected — that is Googlebot-Image."},{"slug":"googleother-video","name":"GoogleOther-Video","operator":"Google","category":"ai-training","category_label":"AI training crawlers","ua":"GoogleOther-Video","robots_token":"GoogleOther-Video","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googleother-video.html","what_it_is":"The video variant of GoogleOther, used for internal Google fetches that do not belong to Search.","cost_of_blocking":"No effect on Search or on Google Video search. Blocks internal Google research fetches of your video files."},{"slug":"googleproducer","name":"GoogleProducer","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"GoogleProducer","robots_token":"GoogleProducer","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googleproducer.html","what_it_is":"Google Publisher Center: fetches the feeds a publisher explicitly supplied for Google News landing pages. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"Your own Google News landing pages stop updating. Only publishers who configured Publisher Center are affected."},{"slug":"gptbot","name":"GPTBot","operator":"OpenAI","category":"ai-training","category_label":"AI training crawlers","ua":"GPTBot","robots_token":"GPTBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://platform.openai.com/docs/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/openai-gptbot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/gptbot.html","what_it_is":"OpenAI's bulk crawler. Pages it fetches may be used to train future OpenAI foundation models. It is not the bot that puts you in ChatGPT's search results, and blocking it does not remove you from them.","cost_of_blocking":"Your content is excluded from training data for future OpenAI models. No effect on ChatGPT search visibility, on citations, or on links a user pastes into ChatGPT."},{"slug":"ia-archiver","name":"ia_archiver","operator":"Internet Archive","category":"archive","category_label":"Archivers","ua":"ia_archiver","robots_token":"ia_archiver","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://archive.org/details/archive.org_bot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ia-archiver.html","what_it_is":"The legacy Alexa/Internet Archive token, still present in most robots.txt files and still occasionally honoured.","cost_of_blocking":"Negligible today; retain for tidiness."},{"slug":"icc-crawler","name":"ICC-Crawler","operator":"NICT","category":"ai-training","category_label":"AI training crawlers","ua":"ICC-Crawler","robots_token":"ICC-Crawler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.nict.go.jp/en/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/icc-crawler.html","what_it_is":"Operated by NICT, Japan's national information and communications research institute. The collected data supports AI research and, per the operator, is also provided to third parties including commercial companies.","cost_of_blocking":"You are excluded from a national research corpus and from the commercial redistributions of it. This is a dataset-shaped block: one refusal, many downstream effects."},{"slug":"imagesiftbot","name":"ImagesiftBot","operator":"Hive AI","category":"dataset","category_label":"Corpus and dataset builders","ua":"ImagesiftBot","robots_token":"ImagesiftBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://imagesift.com/about","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/imagesiftbot.html","what_it_is":"Crawls images for Hive AI's reverse-image and dataset products. Image-heavy sites see this one long before they see the text crawlers.","cost_of_blocking":"Your images stop entering an image dataset and reverse-image index."},{"slug":"img2dataset","name":"img2dataset","operator":"LAION / img2dataset","category":"dataset","category_label":"Corpus and dataset builders","ua":"img2dataset","robots_token":"img2dataset","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://github.com/rom1504/img2dataset","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/img2dataset.html","what_it_is":"The tool used to turn image-URL lists such as LAION's into downloaded training sets. It is run by whoever is building a dataset, not by a single operator.","cost_of_blocking":"Your images are skipped when someone materialises an image-text dataset that references them."},{"slug":"isscyberriskcrawler","name":"ISSCyberRiskCrawler","operator":"ISS Corporate Solutions","category":"ai-training","category_label":"AI training crawlers","ua":"ISSCyberRiskCrawler","robots_token":"ISSCyberRiskCrawler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"disputed","docs":"https://iss-cyber.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/isscyberriskcrawler.html","what_it_is":"Crawls in order to train models that score a company's cyber risk. The ai.robots.txt dataset records the operator as not respecting robots.txt; ISS publishes no compliance statement of its own.","cost_of_blocking":"A rule here is a statement of intent. Your organisation's public footprint still gets scored — by a model trained on everybody else."},{"slug":"kagibot","name":"Kagibot","operator":"Kagi","category":"search","category_label":"Search engines","ua":"Kagibot","robots_token":"Kagibot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://kagi.com/bot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/kagibot.html","what_it_is":"The crawler for Kagi, a paid, ad-free search engine with its own index and its own assistant.","cost_of_blocking":"You leave Kagi's index. Kagi's users are paying to search and skew technical; per visitor this is an expensive block."},{"slug":"klaviyoaibot","name":"KlaviyoAIBot","operator":"Klaviyo","category":"ai-search","category_label":"AI search crawlers","ua":"KlaviyoAIBot","robots_token":"KlaviyoAIBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://help.klaviyo.com/hc/en-us/articles/40496146232219","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/klaviyoaibot.html","what_it_is":"Fetches pages from domains a Klaviyo customer has explicitly connected to their own account, to power Klaviyo's Kai customer agent. It is scoped to connected domains rather than the open web.","cost_of_blocking":"If the connected domain is yours, blocking this breaks the agent you configured. If it is not, this bot should not be reaching you at all."},{"slug":"laiondownloader","name":"LAIONDownloader","operator":"LAION / img2dataset","category":"dataset","category_label":"Corpus and dataset builders","ua":"LAIONDownloader","robots_token":"LAIONDownloader","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"by-design-no","docs":"https://laion.ai/faq/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/laiondownloader.html","what_it_is":"LAION's downloader, used to materialise the image and text datasets the non-profit publishes for machine-learning research. LAION's own FAQ is the source for its robots.txt position.","cost_of_blocking":"Your media is skipped when an open research dataset is built from URL lists. Once a dataset is published, a later block does not remove you from it."},{"slug":"lightpanda","name":"Lightpanda","operator":"Lightpanda","category":"tool","category_label":"Tools and frameworks","ua":"Lightpanda","robots_token":"Lightpanda","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://lightpanda.io/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/lightpanda.html","what_it_is":"A purpose-built headless browser for AI and automation — a runtime, not an operator. Whether robots.txt is honoured is left to whoever runs it, which is what its maintainers say themselves.","cost_of_blocking":"You block a browser, not a company: the same rule stops a scraper and a legitimate automation a customer of yours is running."},{"slug":"linguee-bot","name":"Linguee Bot","operator":"Linguee","category":"ai-training","category_label":"AI training crawlers","ua":"Linguee Bot","robots_token":"Linguee Bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"disputed","docs":"https://www.linguee.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/linguee-bot.html","what_it_is":"Gathers bilingual text for Linguee's translation corpus and the machine translation trained on it. Recorded in the ai.robots.txt dataset as not respecting robots.txt.","cost_of_blocking":"Multilingual pages stop feeding a translation corpus. If your site is translated, being in it is usually a benefit."},{"slug":"mediapartners-google","name":"Mediapartners-Google","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"Mediapartners-Google","robots_token":"Mediapartners-Google","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"own-token-only","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/mediapartners-google.html","what_it_is":"The AdSense crawler. It reads a page so AdSense can choose relevant ads for it, and it is a special-case crawler that ignores the robots.txt * group.","cost_of_blocking":"Pages it cannot read get generic, lower-value AdSense ads or none at all. This is the one block on this list that costs you money directly if you run AdSense."},{"slug":"meta-externalagent","name":"meta-externalagent","operator":"Meta","category":"ai-training","category_label":"AI training crawlers","ua":"meta-externalagent","robots_token":"meta-externalagent","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/meta-externalagent.html","what_it_is":"Meta's AI crawler, gathering training data for Llama and Meta AI. It replaced the older FacebookBot name for this purpose.","cost_of_blocking":"Excluded from Meta AI training. Link previews on Facebook, Instagram and WhatsApp are unaffected — those are a different bot."},{"slug":"meta-externalfetcher","name":"meta-externalfetcher","operator":"Meta","category":"user-fetch","category_label":"User-triggered fetchers","ua":"meta-externalfetcher","robots_token":"meta-externalfetcher","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/meta-externalfetcher.html","what_it_is":"Fetches a page when a Meta AI user asks about a specific link.","cost_of_blocking":"Meta AI cannot read pages users hand it."},{"slug":"meta-webindexer","name":"Meta-WebIndexer","operator":"Meta","category":"ai-search","category_label":"AI search crawlers","ua":"Meta-WebIndexer","robots_token":"Meta-WebIndexer","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/meta-webindexer.html","what_it_is":"Per Meta's crawler documentation, Meta-WebIndexer navigates the web to improve the quality of Meta AI's search results. It is a third Meta token alongside Meta-ExternalAgent and Meta-ExternalFetcher, and the newest of them.","cost_of_blocking":"You leave the index Meta AI answers from across Facebook, Instagram and WhatsApp — the largest assistant install base there is. A robots.txt that names the two older Meta tokens does not cover this one."},{"slug":"mistralai-user","name":"MistralAI-User","operator":"Mistral AI","category":"user-fetch","category_label":"User-triggered fetchers","ua":"MistralAI-User","robots_token":"MistralAI-User","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://docs.mistral.ai/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/mistralai-user.html","what_it_is":"Fetches a page when a Le Chat user asks Mistral's assistant to read it.","cost_of_blocking":"Le Chat cannot open links your readers give it."},{"slug":"mj12bot","name":"MJ12bot","operator":"Majestic","category":"seo","category_label":"SEO and backlink crawlers","ua":"MJ12bot","robots_token":"MJ12bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://mj12bot.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/mj12bot.html","what_it_is":"Majestic's link-graph crawler, run as a distributed community project. Majestic states plainly that it cannot restrict the bot to a fixed set of addresses, and offers a pre-arranged ident string in the request headers instead.","cost_of_blocking":"You leave the Majestic backlink index. No search or AI effect. It supports Crawl-delay, which is usually the better answer than a block."},{"slug":"mojeekbot","name":"MojeekBot","operator":"Mojeek","category":"search","category_label":"Search engines","ua":"MojeekBot","robots_token":"MojeekBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.mojeek.com/bot.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/mojeekbot.html","what_it_is":"Mojeek's crawler. Mojeek runs one of the few genuinely independent web indexes — not a front end over Bing or Google — so it is one of the few blocks that removes you from an index nobody else can put you back into. Its documentation states it obeys the first record whose User-Agent contains MojeekBot, falling back to *.","cost_of_blocking":"You leave an independent index that other privacy-focused search products draw on. Small traffic, disproportionate long-term cost to web plurality."},{"slug":"oai-searchbot","name":"OAI-SearchBot","operator":"OpenAI","category":"ai-search","category_label":"AI search crawlers","ua":"OAI-SearchBot","robots_token":"OAI-SearchBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://platform.openai.com/docs/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/openai-searchbot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/oai-searchbot.html","what_it_is":"Builds the index ChatGPT search answers from. Content it collects is used for retrieval and citation, not for model training.","cost_of_blocking":"High. Blocking this removes you from ChatGPT search results and from the source links ChatGPT shows. This is the single most expensive block on this list for anyone who wants to be cited by an assistant."},{"slug":"omgili","name":"omgili","operator":"Webz.io","category":"dataset","category_label":"Corpus and dataset builders","ua":"omgili","robots_token":"omgili","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://webz.io/blog/machine-learning/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/omgili.html","what_it_is":"The older robots token for the same Webz.io collection, still honoured and still worth listing.","cost_of_blocking":"Same as omgilibot."},{"slug":"omgilibot","name":"omgilibot","operator":"Webz.io","category":"dataset","category_label":"Corpus and dataset builders","ua":"omgilibot","robots_token":"omgilibot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://webz.io/blog/machine-learning/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/omgilibot.html","what_it_is":"Webz.io's crawler, collecting web and forum text sold as datasets, including to model builders.","cost_of_blocking":"Exclusion from a commercial dataset resold to third parties."},{"slug":"panscient","name":"Panscient","operator":"Panscient","category":"dataset","category_label":"Corpus and dataset builders","ua":"panscient.com","robots_token":"panscient.com","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://panscient.com/faq.htm","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/panscient.html","what_it_is":"Compiles structured data about businesses and business professionals using machine learning. Panscient's FAQ states it obeys robots.txt.","cost_of_blocking":"Your company pages stop feeding a business-data product. No effect on search or assistants."},{"slug":"perplexity-user","name":"Perplexity-User","operator":"Perplexity","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Perplexity-User","robots_token":"Perplexity-User","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://docs.perplexity.ai/guides/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/perplexity-user.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/perplexity-user.html","what_it_is":"Fetches a page because a Perplexity user asked for it. Perplexity documents that this fetch is user-initiated and is therefore not governed by robots.txt — a robots rule will not stop it, by stated policy.","cost_of_blocking":"Not controllable via robots.txt. If you must stop it, verify by the published IP ranges and block at the edge — and accept that users who ask for your page get an error."},{"slug":"perplexitybot","name":"PerplexityBot","operator":"Perplexity","category":"ai-search","category_label":"AI search crawlers","ua":"PerplexityBot","robots_token":"PerplexityBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://docs.perplexity.ai/guides/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/perplexity-bot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/perplexitybot.html","what_it_is":"Builds Perplexity's search index. Perplexity is citation-heavy by product design, so inclusion here converts to referral traffic more directly than most AI surfaces.","cost_of_blocking":"You stop being indexed and cited by Perplexity, and lose the referral clicks its citations produce."},{"slug":"petalbot","name":"PetalBot","operator":"Huawei","category":"search","category_label":"Search engines","ua":"PetalBot","robots_token":"PetalBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://aspiegel.com/petalbot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/petalbot.html","what_it_is":"Huawei's crawler for Petal Search, shipped as the default search on Huawei devices.","cost_of_blocking":"Removal from Petal Search. Frequently blocked for volume rather than for policy."},{"slug":"phindbot","name":"PhindBot","operator":"Phind","category":"ai-search","category_label":"AI search crawlers","ua":"PhindBot","robots_token":"PhindBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.phind.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/phindbot.html","what_it_is":"Phind is an answer engine for developers that combines live web search with its own models. This is the crawler behind those answers.","cost_of_blocking":"You stop being cited in answers to technical questions — which, for documentation and reference sites, is the exact audience most worth keeping."},{"slug":"pinterestbot","name":"Pinterestbot","operator":"Pinterest","category":"search","category_label":"Search engines","ua":"Pinterestbot","robots_token":"Pinterestbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://help.pinterest.com/en/business/article/pinterest-crawler","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/pinterestbot.html","what_it_is":"Pinterest's crawler. It indexes pages so people can find them on Pinterest and re-reads product pages to keep price and title on a Pin current. Pinterest states that content it crawls is not used to train their Canvas image generation model.","cost_of_blocking":"Pins pointing at your site go stale — wrong prices, dead links — and new content stops being indexed. For a retailer this is one of the more expensive blocks on the list."},{"slug":"poseidon-research-crawler","name":"Poseidon Research Crawler","operator":"Poseidon Research","category":"ai-training","category_label":"AI training crawlers","ua":"Poseidon Research Crawler","robots_token":"Poseidon Research Crawler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.poseidonresearch.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/poseidon-research-crawler.html","what_it_is":"A crawler run by Poseidon Research, a lab working on interpretability research for AI systems.","cost_of_blocking":"Exclusion from an interpretability research corpus. No published compliance statement."},{"slug":"qualifiedbot","name":"QualifiedBot","operator":"Qualified","category":"ai-search","category_label":"AI search crawlers","ua":"QualifiedBot","robots_token":"QualifiedBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.qualified.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/qualifiedbot.html","what_it_is":"Analyses a customer's website so Qualified's AI sales chatbots can answer questions about it in context.","cost_of_blocking":"A chatbot on a site that licensed the product loses context. If that site is yours, this block is self-inflicted."},{"slug":"quillbot","name":"QuillBot","operator":"QuillBot","category":"ai-training","category_label":"AI training crawlers","ua":"QuillBot","robots_token":"QuillBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://quillbot.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/quillbot.html","what_it_is":"Operated by QuillBot as part of its writing, paraphrasing and AI-detection products. The dataset also records a second token, quillbot.com, for the same operator.","cost_of_blocking":"Exclusion from QuillBot's corpus. No compliance statement is published, so the rule is a request."},{"slug":"qwantbot","name":"Qwantbot","operator":"Qwant","category":"search","category_label":"Search engines","ua":"Qwantbot","robots_token":"Qwantbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://help.qwant.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/qwantbot.html","what_it_is":"Qwant's crawler. Qwant documents that the string Qwantbot always appears in its user-agents whatever the crawler version, which is what makes a substring match safe here.","cost_of_blocking":"You leave the index behind Qwant, the French privacy-focused engine, and the products that federate it."},{"slug":"qwantbot-news","name":"Qwantbot-news","operator":"Qwant","category":"search","category_label":"Search engines","ua":"Qwantbot-news","robots_token":"Qwantbot-news","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://help.qwant.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/qwantbot-news.html","what_it_is":"The news variant of Qwant's crawler, documented alongside the main one and carrying the same Qwantbot substring.","cost_of_blocking":"Your articles stop appearing in Qwant News. A rule for Qwantbot as a substring already catches both."},{"slug":"reflectionbot","name":"Reflectionbot","operator":"Reflection AI","category":"ai-training","category_label":"AI training crawlers","ua":"Reflectionbot","robots_token":"Reflectionbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://reflection.ai/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/reflectionbot.html","what_it_is":"An undocumented crawler whose user-agent links to Reflection AI, a company building AI models. The link in the user-agent is the only public statement of purpose that exists.","cost_of_blocking":"Unknown by construction — which is itself the reason some people block it. Nothing user-facing depends on it."},{"slug":"rogerbot","name":"rogerbot","operator":"Moz","category":"seo","category_label":"SEO and backlink crawlers","ua":"rogerbot","robots_token":"rogerbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://moz.com/help/moz-procedures/crawlers/rogerbot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/rogerbot.html","what_it_is":"Moz's Campaign crawler, which audits a site its own owner registered. Moz states there is no IP range for it — identification is by user-agent only.","cost_of_blocking":"Moz Pro site audits of your domain stop. If the domain is yours and you use Moz, blocking this breaks your own reports."},{"slug":"sbintuitionsbot","name":"SBIntuitionsBot","operator":"SB Intuitions","category":"ai-training","category_label":"AI training crawlers","ua":"SBIntuitionsBot","robots_token":"SBIntuitionsBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.sbintuitions.co.jp/en/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/sbintuitionsbot.html","what_it_is":"SB Intuitions is SoftBank's Japanese LLM lab; this crawler gathers data used in that model development and in information analysis. The operator publishes a dedicated bot page.","cost_of_blocking":"Your content is excluded from a Japanese-language foundation-model corpus. Nothing user-facing changes."},{"slug":"scrapy","name":"Scrapy","operator":"Scrapy project","category":"tool","category_label":"Tools and frameworks","ua":"Scrapy","robots_token":"Scrapy","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://scrapy.org/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/scrapy.html","what_it_is":"Not an operator: the default user-agent of the most common Python crawling framework. Anyone can be behind it. Modern Scrapy obeys robots.txt by default, which is why the default UA is still worth a rule.","cost_of_blocking":"You block a very large tail of unattributed one-off crawlers, and also every well-behaved researcher who did not change the default."},{"slug":"screaming-frog-seo-spider","name":"Screaming Frog SEO Spider","operator":"Screaming Frog","category":"tool","category_label":"Tools and frameworks","ua":"Screaming Frog SEO Spider","robots_token":"Screaming Frog SEO Spider","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.screamingfrog.co.uk/seo-spider/user-agent/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/screaming-frog-seo-spider.html","what_it_is":"Not an operator: desktop crawling software that anybody can point at any site. The default user-agent identifies the tool, not who is running it, and the operator of the moment is whoever pressed start.","cost_of_blocking":"You block a consultant auditing your own site as often as you block a stranger. Treat it as a rate-limit question, not a consent one — and note that the user-agent is configurable, so a block is advisory."},{"slug":"semrushbot","name":"SemrushBot","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot","robots_token":"SemrushBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot.html","what_it_is":"Semrush's backlink and keyword crawler. It is not an AI crawler, but it is usually in the top three by volume on any site, and it is the cheapest block on this list.","cost_of_blocking":"Your competitors' Semrush reports get thinner, and so do yours. No user-facing effect."},{"slug":"semrushbot-ba","name":"SemrushBot-BA","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot-BA","robots_token":"SemrushBot-BA","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-ba.html","what_it_is":"The Backlink Audit crawler. It re-checks links pointing at a customer's site, which means it lands on the sites doing the linking.","cost_of_blocking":"None to you. It costs the site being audited a little accuracy in their backlink report."},{"slug":"semrushbot-esi","name":"SemrushBot-ESI","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot-ESI","robots_token":"SemrushBot-ESI","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-esi.html","what_it_is":"The crawler for Semrush Enterprise Site Intelligence, the enterprise tier's own site analysis.","cost_of_blocking":"Enterprise customers lose analysis of your domain. Nothing user-facing."},{"slug":"semrushbot-ft","name":"SemrushBot-FT","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot-FT","robots_token":"SemrushBot-FT","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-ft.html","what_it_is":"Fetches full text for the Plagiarism Checker and similar text-comparison tools.","cost_of_blocking":"Your text stops being compared against other people's submissions — which also means copies of your text are less likely to be caught."},{"slug":"semrushbot-ocob","name":"SemrushBot-OCOB","operator":"Semrush","category":"ai-training","category_label":"AI training crawlers","ua":"SemrushBot-OCOB","robots_token":"SemrushBot-OCOB","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-ocob.html","what_it_is":"Semrush's separately-tokenised crawler for its AI content tooling, split out so SEO crawling and AI reuse can be answered differently.","cost_of_blocking":"Exclusion from Semrush's AI corpus, with its SEO crawl unaffected."},{"slug":"semrushbot-si","name":"SemrushBot-SI","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot-SI","robots_token":"SemrushBot-SI","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-si.html","what_it_is":"Fetches pages for the On Page SEO Checker and similar advisory tools.","cost_of_blocking":"Nothing user-facing. Semrush customers lose on-page suggestions for pages on your domain."},{"slug":"semrushbot-swa","name":"SemrushBot-SWA","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot-SWA","robots_token":"SemrushBot-SWA","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-swa.html","what_it_is":"Checks whether a URL is reachable, for the SEO Writing Assistant. One request per URL a writer references, not a crawl.","cost_of_blocking":"Writers using Semrush's assistant see your links reported as unreachable."},{"slug":"seokicks","name":"SEOkicks","operator":"SEOkicks","category":"seo","category_label":"SEO and backlink crawlers","ua":"SEOkicks","robots_token":"SEOkicks","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.seokicks.de/robot.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/seokicks.html","what_it_is":"A German backlink index. Its documentation names SEOkicks as the user-agent to use in robots.txt.","cost_of_blocking":"You leave a regional backlink index. Nothing else changes."},{"slug":"serpstatbot","name":"serpstatbot","operator":"Serpstat","category":"seo","category_label":"SEO and backlink crawlers","ua":"serpstatbot","robots_token":"serpstatbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://serpstatbot.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/serpstatbot.html","what_it_is":"Serpstat's backlink crawler. It documents support for Crawl-delay up to 20 seconds, including a delay set on the * group.","cost_of_blocking":"You leave Serpstat's link index. Try Crawl-delay first — they honour it, and a slow crawler is cheaper to keep than to fight."},{"slug":"seznambot","name":"SeznamBot","operator":"Seznam","category":"search","category_label":"Search engines","ua":"SeznamBot","robots_token":"SeznamBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://napoveda.seznam.cz/en/seznamzbozi/subject-matter-crawler/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/seznambot.html","what_it_is":"Seznam's crawler — the dominant search engine in the Czech Republic and one of the few national engines with its own index.","cost_of_blocking":"Removal from Seznam. Also removes you from its IndexNow endpoint's usefulness."},{"slug":"shapbot","name":"ShapBot","operator":"Parallel","category":"ai-search","category_label":"AI search crawlers","ua":"ShapBot","robots_token":"ShapBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://docs.parallel.ai/features/crawler","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/shapbot.html","what_it_is":"Parallel's crawler. It collects and structures web content to power the search, extraction and deep-research APIs that Parallel sells to agent builders.","cost_of_blocking":"Agents using Parallel's research API lose you as a source. Parallel documents robots.txt compliance, so a rule works."},{"slug":"sidetrade-indexer-bot","name":"Sidetrade indexer bot","operator":"Sidetrade","category":"ai-training","category_label":"AI training crawlers","ua":"Sidetrade indexer bot","robots_token":"Sidetrade indexer bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.sidetrade.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/sidetrade-indexer-bot.html","what_it_is":"Sidetrade extracts web data for a range of uses including training its AI products for order-to-cash and customer-data work.","cost_of_blocking":"Exclusion from a commercial B2B dataset. The operator publishes no robots.txt statement, so treat the rule as a request rather than a control."},{"slug":"siteauditbot","name":"SiteAuditBot","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SiteAuditBot","robots_token":"SiteAuditBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/siteauditbot.html","what_it_is":"Semrush's Site Audit crawler: it walks a site a customer owns and reports technical SEO problems. Semrush names it as the token to block for that product.","cost_of_blocking":"Site Audit reports on your own domain stop working. If a customer is auditing your site with your permission, blocking this breaks their tooling and nothing of yours."},{"slug":"slackbot","name":"Slackbot","operator":"Slack","category":"preview","category_label":"Link preview fetchers","ua":"Slackbot","robots_token":"Slackbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://api.slack.com/robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/slackbot.html","what_it_is":"The other half of Slack's pair: the agent that reads robots.txt and handles Slack's non-unfurl fetches. Slack documents both strings on one page.","cost_of_blocking":"Slack stops being able to read your robots.txt, which is a strange thing to want. Block the link expander instead if that is the goal."},{"slug":"slackbot-linkexpanding","name":"Slackbot-LinkExpanding","operator":"Slack","category":"preview","category_label":"Link preview fetchers","ua":"Slackbot-LinkExpanding","robots_token":"Slackbot-LinkExpanding","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://api.slack.com/robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/slackbot-linkexpanding.html","what_it_is":"Fetches a page to build the unfurl card when somebody pastes your link into Slack. One paste, one fetch.","cost_of_blocking":"Your links appear in Slack as bare URLs. Inside working teams that quietly costs you clicks, and nothing is gained."},{"slug":"splitsignalbot","name":"SplitSignalBot","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SplitSignalBot","robots_token":"SplitSignalBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/splitsignalbot.html","what_it_is":"Runs SEO A/B tests on a customer's own site with the SplitSignal tool.","cost_of_blocking":"A site owner's own A/B testing stops. Only relevant on domains whose owner uses the product."},{"slug":"storebot-google","name":"Storebot-Google","operator":"Google","category":"search","category_label":"Search engines","ua":"Storebot-Google","robots_token":"Storebot-Google","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/storebot-google.html","what_it_is":"Checks shopping and checkout flows for Google's shopping surfaces.","cost_of_blocking":"Product listings may lose shopping-specific enrichment. Irrelevant to non-commerce sites."},{"slug":"terracotta","name":"TerraCotta","operator":"Ceramic AI","category":"ai-search","category_label":"AI search crawlers","ua":"TerraCotta","robots_token":"TerraCotta","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://ceramic.ai/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/terracotta.html","what_it_is":"Ceramic AI's crawler, which indexes public content for a web-scale search API aimed at LLMs and agents.","cost_of_blocking":"You are absent from another agent-facing retrieval index. Ceramic documents that it obeys robots.txt."},{"slug":"thinkbot","name":"Thinkbot","operator":"Thinkbot","category":"dataset","category_label":"Corpus and dataset builders","ua":"Thinkbot","robots_token":"Thinkbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"disputed","docs":"https://www.thinkbot.agency","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/thinkbot.html","what_it_is":"Collects pages for analysis of how sites are adopting AI and automation. The ai.robots.txt dataset records the operator as not respecting robots.txt.","cost_of_blocking":"Exclusion from a market-research dataset. Expect to enforce this at the edge rather than in robots.txt."},{"slug":"tiktokspider","name":"TikTokSpider","operator":"ByteDance","category":"ai-training","category_label":"AI training crawlers","ua":"TikTokSpider","robots_token":"TikTokSpider","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"disputed","docs":"https://www.bytespider.net/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/tiktokspider.html","what_it_is":"A second ByteDance crawler identifying with TikTok, collecting page content for the same family of models.","cost_of_blocking":"Little to lose unless TikTok search referral matters to you."},{"slug":"timpibot","name":"Timpibot","operator":"Timpi","category":"search","category_label":"Search engines","ua":"Timpibot","robots_token":"Timpibot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://timpi.io/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/timpibot.html","what_it_is":"A distributed crawler building an independent search index outside the Google/Bing duopoly.","cost_of_blocking":"Absence from a small independent index."},{"slug":"velenpublicwebcrawler","name":"VelenPublicWebCrawler","operator":"Hunter (Velen)","category":"dataset","category_label":"Corpus and dataset builders","ua":"VelenPublicWebCrawler","robots_token":"VelenPublicWebCrawler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://velen.io/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/velenpublicwebcrawler.html","what_it_is":"Hunter's crawler, written in Go, building business datasets and machine-learning models from public pages. Its page states it follows robots.txt and meta directives and never fetches more than one page every two seconds.","cost_of_blocking":"Your company pages stop feeding a B2B contact and company dataset. The crawl rate it documents makes this one of the cheapest visitors to simply allow."},{"slug":"webzio-extended","name":"Webzio-Extended","operator":"Webz.io","category":"ai-training","category_label":"AI training crawlers","ua":"Webzio-Extended","robots_token":"Webzio-Extended","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://webz.io/blog/machine-learning/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/webzio-extended.html","what_it_is":"Webz.io's opt-out token specifically for AI training reuse, in the pattern Google and Apple established.","cost_of_blocking":"Your content is excluded from the AI-training tier of Webz.io's product while ordinary collection continues."},{"slug":"wpbot","name":"wpbot","operator":"QuantumCloud","category":"tool","category_label":"Tools and frameworks","ua":"wpbot","robots_token":"wpbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.quantumcloud.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/wpbot.html","what_it_is":"Supports the AI Chatbot for WordPress plugin: it reads pages so the plugin can answer from a site's own content. The operator provides an opt-out through a form rather than through robots.txt.","cost_of_blocking":"A WordPress site running that plugin loses its own content as an answer source. Only relevant where the plugin is installed."},{"slug":"yak","name":"YaK","operator":"Meltwater","category":"dataset","category_label":"Corpus and dataset builders","ua":"YaK","robots_token":"YaK","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.meltwater.com/en/suite/consumer-intelligence","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yak.html","what_it_is":"Meltwater's crawler, feeding the live data stream behind its media-monitoring and consumer-intelligence suite.","cost_of_blocking":"Your content stops appearing in Meltwater's media monitoring — which is how PR teams find out you were mentioned. Some publishers want to be in it."},{"slug":"yandexadditional","name":"YandexAdditional","operator":"Yandex","category":"ai-training","category_label":"AI training crawlers","ua":"YandexAdditional","robots_token":"YandexAdditional","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexadditional.html","what_it_is":"The token that controls whether already-indexed pages may appear in Search with Yandex AI answers. Yandex's table says it makes no indexing requests of its own — it exists so a site can opt out of the generative answer without leaving the index.","cost_of_blocking":"You disappear from Yandex's AI answers while staying in Yandex Search. This is Yandex's equivalent of Google-Extended, and it is the cheap opt-out most people are looking for."},{"slug":"yandexadditionalbot","name":"YandexAdditionalBot","operator":"Yandex","category":"ai-training","category_label":"AI training crawlers","ua":"YandexAdditionalBot","robots_token":"YandexAdditionalBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexadditionalbot.html","what_it_is":"The second token Yandex publishes for the same AI-answers opt-out. Both names appear in Yandex's own robot list, so a robots.txt that names only one of them is half a policy.","cost_of_blocking":"Same as YandexAdditional: out of Yandex's AI answers, still in Yandex Search. Name both tokens or neither."},{"slug":"yandexblogs","name":"YandexBlogs","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexBlogs","robots_token":"YandexBlogs","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexblogs.html","what_it_is":"Yandex's blog-search robot; it indexes post comments as well as posts.","cost_of_blocking":"Comment threads and blog posts stop being findable through Yandex blog search."},{"slug":"yandexbot","name":"YandexBot","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexBot","robots_token":"YandexBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/robot-workings/check-yandex-robots.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexbot.html","what_it_is":"Yandex's search crawler, which also feeds Alice and Yandex's generative answers.","cost_of_blocking":"Removal from Yandex Search. Verify with reverse DNS to a yandex.ru, yandex.net or yandex.com host — YandexBot is among the most-spoofed user-agents there is."},{"slug":"yandexcalendar","name":"YandexCalendar","operator":"Yandex","category":"user-fetch","category_label":"User-triggered fetchers","ua":"YandexCalendar","robots_token":"YandexCalendar","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexcalendar.html","what_it_is":"Downloads calendar files a user subscribed to. Yandex notes these files are often in directories that are disallowed for indexing, which is why the general rules are not applied.","cost_of_blocking":"Users who subscribed to a calendar you publish stop receiving updates."},{"slug":"yandexcombot","name":"YandexComBot","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexComBot","robots_token":"YandexComBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexcombot.html","what_it_is":"Indexes content for Yandex search in languages other than Russian. Yandex documents that it can index content when there is no explicit robot-specific restriction — a * group is not one.","cost_of_blocking":"You leave Yandex's non-Russian index. A rule naming this token is the only one that works on it."},{"slug":"yandexdirect","name":"YandexDirect","operator":"Yandex","category":"tool","category_label":"Tools and frameworks","ua":"YandexDirect","robots_token":"YandexDirect","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexdirect.html","what_it_is":"Reads the content of Yandex Advertising Network partner pages to work out their topic so relevant ads can be matched. Documented as not taking the general robots.txt rules into account.","cost_of_blocking":"Ads on your pages become less relevant and earn less. Relevant only if you monetise with Yandex's network."},{"slug":"yandexfavicons","name":"YandexFavicons","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexFavicons","robots_token":"YandexFavicons","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexfavicons.html","what_it_is":"Downloads your favicon so Yandex can show it beside your result. Documented as not taking the general robots.txt rules into account.","cost_of_blocking":"Your results in Yandex lose their icon. Cosmetic, and a * rule will not achieve it anyway."},{"slug":"yandeximages","name":"YandexImages","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexImages","robots_token":"YandexImages","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandeximages.html","what_it_is":"Indexes images for Yandex Images. Yandex's robot table marks it as taking the general robots.txt rules into account.","cost_of_blocking":"Your images leave Yandex Images, which is a large share of image search in Russian-speaking markets."},{"slug":"yandexmarket","name":"YandexMarket","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexMarket","robots_token":"YandexMarket","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexmarket.html","what_it_is":"The robot behind Yandex Market, Yandex's shopping comparison service. Version 1.0 is documented as obeying the general rules; version 2.0 is documented as not.","cost_of_blocking":"Your products stop being listed and priced in Yandex Market. For a retailer in that market this is a revenue block, not a bandwidth one."},{"slug":"yandexmedia","name":"YandexMedia","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexMedia","robots_token":"YandexMedia","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexmedia.html","what_it_is":"Indexes multimedia data for Yandex. Takes the general robots.txt rules into account.","cost_of_blocking":"Your multimedia content stops appearing in Yandex's media surfaces."},{"slug":"yandexmetrika","name":"YandexMetrika","operator":"Yandex","category":"tool","category_label":"Tools and frameworks","ua":"YandexMetrika","robots_token":"YandexMetrika","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"by-design-no","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexmetrika.html","what_it_is":"Yandex Metrica's own fetcher. Two of its versions — the 2.0 yabs01 availability checker and the 4.0 CSS cache for Webvisor — are documented in Yandex's table as not using robots.txt at all.","cost_of_blocking":"Nothing you can enforce through robots.txt. If you run Metrica, its session replay loses your stylesheets and renders your pages wrong."},{"slug":"yandexmobilebot","name":"YandexMobileBot","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexMobileBot","robots_token":"YandexMobileBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexmobilebot.html","what_it_is":"Decides whether a page's layout is suitable for mobile devices. Yandex's table marks it as NOT taking the general robots.txt rules into account, so a * group does not stop it — a group named YandexMobileBot does.","cost_of_blocking":"Yandex loses its mobile-friendliness signal for your pages, which affects how they are ranked and rendered on phones."},{"slug":"yandexrenderresourcesbot","name":"YandexRenderResourcesBot","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexRenderResourcesBot","robots_token":"YandexRenderResourcesBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexrenderresourcesbot.html","what_it_is":"Loads the CSS, JavaScript and images Yandex needs to render a page. Yandex documents the exact rule: it ignores robots.txt for a resource when the HTML page using it is allowed, and does not fetch the resource when that page is disallowed.","cost_of_blocking":"Yandex renders your pages without their stylesheets or scripts and ranks what it sees. This is the classic accidental self-inflicted ranking loss."},{"slug":"yandexscreenshotbot","name":"YandexScreenshotBot","operator":"Yandex","category":"tool","category_label":"Tools and frameworks","ua":"YandexScreenshotBot","robots_token":"YandexScreenshotBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexscreenshotbot.html","what_it_is":"Takes a screenshot of a page. Documented as not taking the general robots.txt rules into account.","cost_of_blocking":"Yandex surfaces that show a page thumbnail show nothing for you."},{"slug":"yandexvideo","name":"YandexVideo","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexVideo","robots_token":"YandexVideo","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexvideo.html","what_it_is":"Indexes video for Yandex video search. Obeys the general robots.txt rules per Yandex's own table.","cost_of_blocking":"Removal from Yandex video search. Note that a second robot, YandexVideoParser, does the same job and is documented as NOT taking the general rules into account."},{"slug":"yandexwebmaster","name":"YandexWebmaster","operator":"Yandex","category":"tool","category_label":"Tools and frameworks","ua":"YandexWebmaster","robots_token":"YandexWebmaster","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexwebmaster.html","what_it_is":"The fetcher behind Yandex Webmaster, the console a site owner uses to inspect their own site.","cost_of_blocking":"Your own Yandex Webmaster checks stop working. Blocking this only hurts you."},{"slug":"yeti","name":"Yeti","operator":"Naver","category":"search","category_label":"Search engines","ua":"Yeti","robots_token":"Yeti","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://searchadvisor.naver.com/guide/seo-basic-crawl","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yeti.html","what_it_is":"Naver's crawler. Naver is South Korea's largest search portal and runs its own index and its own generative answers.","cost_of_blocking":"Removal from Naver, which is most of Korean search."},{"slug":"youbot","name":"YouBot","operator":"You.com","category":"ai-search","category_label":"AI search crawlers","ua":"YouBot","robots_token":"YouBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://about.you.com/youbot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/youbot.html","what_it_is":"You.com's crawler, feeding its AI search product and its search API.","cost_of_blocking":"Removal from You.com's index and from answers built on its API."}]}
|
|
1
|
+
{"version":"1.3.0","generated_at":"2026-09-05T17:25:28+00:00","source":"AI Crawler Index — https://www.pathwren.workers.dev — independent, non-commercial; data CC0-1.0","data_url":"https://www.pathwren.workers.dev/c/rubygems-registry/data/agents.json","window":"Table reviewed 2026-09-05; every record checked against its operator's own published documentation. IP-range mirrors behind the index refresh every six hours; this bundle is a snapshot, not a live feed.","license":{"code":"MIT","data":"CC0-1.0"},"categories":{"ai-training":{"label":"AI training crawlers","description":"Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today."},"ai-search":{"label":"AI search crawlers","description":"Build the retrieval index an assistant answers and cites from. These are the crawlers that send you traffic; blocking them is the expensive mistake in this space."},"user-fetch":{"label":"User-triggered fetchers","description":"Fetch one page because a person asked for it, right then. One human intent, one request. Blocking them produces a visible error for a real reader."},"dataset":{"label":"Corpus and dataset builders","description":"Crawl the web into a published or resold dataset that other people train on. Highest leverage per block, longest delay before any effect."},"search":{"label":"Search engines","description":"Classic index-and-rank crawlers. Several also feed their operator's generative answers, which is why the AI opt-out for Google and Apple is a token rather than a block."},"seo":{"label":"SEO and backlink crawlers","description":"Commercial link-graph tooling. No user-facing effect either way, and usually a large share of your bot bandwidth."},"archive":{"label":"Archivers","description":"Preservation crawlers. Their output is public and permanent, which makes them a separate decision from the AI one."},"preview":{"label":"Link preview fetchers","description":"Read your Open Graph tags when someone shares a link. Blocking these is almost always an accident."},"tool":{"label":"Tools and frameworks","description":"Not operators: crawling software anyone can run. The party behind the request is unknown, so treat them as a rate-limit question rather than a consent question."}},"ai_categories":["ai-training","ai-search","user-fetch","dataset"],"regex":{"all":"(AdsBot\\-Google|AdsBot\\-Google\\-Mobile|AdsBot\\-Google\\-Mobile\\-Apps|AhrefsBot|AhrefsSiteAudit|AI2Bot|Ai2Bot\\-Dolma|aiHitBot|AIWebIndex|Amazonbot|Andibot|Anomura|anthropic\\-ai|APIs\\-Google|Applebot|archive\\.org_bot|atlassian\\-bot|AwarioRssBot|AwarioSmartBot|Baiduspider|barkrowler|bedrockbot|bingbot|Bytespider|CCBot|ChatGPT\\ Agent|ChatGPT\\-User|Claude\\-SearchBot|Claude\\-User|Claude\\-Web|ClaudeBot|Cloudflare\\-AutoRAG|cohere\\-ai|cohere\\-training\\-data\\-crawler|Cotoyogi|Crawl4AI|Crawlspace|DataForSeoBot|Diffbot|dotbot|DuckAssistBot|DuckDuckBot|EchoboxBot|ExaSearchBot|FacebookBot|facebookexternalhit|Factset_spyderbot|FeedFetcher\\-Google|FirecrawlAgent|Google\\-Agent|Google\\-CloudVertexBot|Google\\-CWS|Google\\-GeminiNotebook|Google\\-InspectionTool|Google\\-Pinpoint|Google\\-Read\\-Aloud|Google\\-Safety|Google\\-Site\\-Verification|Googlebot|Googlebot\\-Image|Googlebot\\-News|Googlebot\\-Video|GoogleMessages|GoogleOther|GoogleOther\\-Image|GoogleOther\\-Video|GoogleProducer|GPTBot|ia_archiver|ICC\\-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|Kagibot|KlaviyoAIBot|LAIONDownloader|Lightpanda|Linguee\\ Bot|Mediapartners\\-Google|meta\\-externalagent|meta\\-externalfetcher|Meta\\-WebIndexer|MistralAI\\-User|MJ12bot|MojeekBot|OAI\\-SearchBot|omgili|omgilibot|panscient\\.com|Perplexity\\-User|PerplexityBot|PetalBot|PhindBot|Pinterestbot|Poseidon\\ Research\\ Crawler|QualifiedBot|QuillBot|Qwantbot|Qwantbot\\-news|Reflectionbot|rogerbot|SBIntuitionsBot|Scrapy|Screaming\\ Frog\\ SEO\\ Spider|SemrushBot|SemrushBot\\-BA|SemrushBot\\-ESI|SemrushBot\\-FT|SemrushBot\\-OCOB|SemrushBot\\-SI|SemrushBot\\-SWA|SEOkicks|serpstatbot|SeznamBot|ShapBot|Sidetrade\\ indexer\\ bot|SiteAuditBot|Slackbot|Slackbot\\-LinkExpanding|SplitSignalBot|Storebot\\-Google|TerraCotta|Thinkbot|TikTokSpider|Timpibot|VelenPublicWebCrawler|Webzio\\-Extended|wpbot|YaK|YandexAdditional|YandexAdditionalBot|YandexBlogs|YandexBot|YandexCalendar|YandexComBot|YandexDirect|YandexFavicons|YandexImages|YandexMarket|YandexMedia|YandexMetrika|YandexMobileBot|YandexRenderResourcesBot|YandexScreenshotBot|YandexVideo|YandexWebmaster|Yeti|YouBot)","ai_only":"(AI2Bot|Ai2Bot\\-Dolma|aiHitBot|AIWebIndex|Amazonbot|Andibot|Anomura|anthropic\\-ai|atlassian\\-bot|AwarioRssBot|AwarioSmartBot|bedrockbot|Bytespider|CCBot|ChatGPT\\ Agent|ChatGPT\\-User|Claude\\-SearchBot|Claude\\-User|Claude\\-Web|ClaudeBot|Cloudflare\\-AutoRAG|cohere\\-ai|cohere\\-training\\-data\\-crawler|Cotoyogi|Diffbot|DuckAssistBot|EchoboxBot|ExaSearchBot|FacebookBot|Factset_spyderbot|Google\\-Agent|Google\\-CloudVertexBot|Google\\-GeminiNotebook|Google\\-Pinpoint|Google\\-Read\\-Aloud|GoogleOther|GoogleOther\\-Image|GoogleOther\\-Video|GPTBot|ICC\\-Crawler|ImagesiftBot|img2dataset|ISSCyberRiskCrawler|KlaviyoAIBot|LAIONDownloader|Linguee\\ Bot|meta\\-externalagent|meta\\-externalfetcher|Meta\\-WebIndexer|MistralAI\\-User|OAI\\-SearchBot|omgili|omgilibot|panscient\\.com|Perplexity\\-User|PerplexityBot|PhindBot|Poseidon\\ Research\\ Crawler|QualifiedBot|QuillBot|Reflectionbot|SBIntuitionsBot|SemrushBot\\-OCOB|ShapBot|Sidetrade\\ indexer\\ bot|TerraCotta|Thinkbot|TikTokSpider|VelenPublicWebCrawler|Webzio\\-Extended|YaK|YandexAdditional|YandexAdditionalBot|YandexCalendar|YouBot)","by_category":{"ai-training":"(anthropic\\-ai|Bytespider|ClaudeBot|cohere\\-training\\-data\\-crawler|Cotoyogi|FacebookBot|Factset_spyderbot|GoogleOther|GoogleOther\\-Image|GoogleOther\\-Video|GPTBot|ICC\\-Crawler|ISSCyberRiskCrawler|Linguee\\ Bot|meta\\-externalagent|Poseidon\\ Research\\ Crawler|QuillBot|Reflectionbot|SBIntuitionsBot|SemrushBot\\-OCOB|Sidetrade\\ indexer\\ bot|TikTokSpider|Webzio\\-Extended|YandexAdditional|YandexAdditionalBot)","ai-search":"(AIWebIndex|Amazonbot|Andibot|Anomura|atlassian\\-bot|bedrockbot|Claude\\-SearchBot|Claude\\-Web|Cloudflare\\-AutoRAG|DuckAssistBot|ExaSearchBot|Google\\-CloudVertexBot|KlaviyoAIBot|Meta\\-WebIndexer|OAI\\-SearchBot|PerplexityBot|PhindBot|QualifiedBot|ShapBot|TerraCotta|YouBot)","user-fetch":"(ChatGPT\\ Agent|ChatGPT\\-User|Claude\\-User|cohere\\-ai|Google\\-Agent|Google\\-GeminiNotebook|Google\\-Pinpoint|Google\\-Read\\-Aloud|meta\\-externalfetcher|MistralAI\\-User|Perplexity\\-User|YandexCalendar)","search":"(Applebot|Baiduspider|bingbot|DuckDuckBot|Googlebot|Googlebot\\-Image|Googlebot\\-News|Googlebot\\-Video|Kagibot|MojeekBot|PetalBot|Pinterestbot|Qwantbot|Qwantbot\\-news|SeznamBot|Storebot\\-Google|Timpibot|YandexBlogs|YandexBot|YandexComBot|YandexFavicons|YandexImages|YandexMarket|YandexMedia|YandexMobileBot|YandexRenderResourcesBot|YandexVideo|Yeti)","tool":"(AdsBot\\-Google|AdsBot\\-Google\\-Mobile|AdsBot\\-Google\\-Mobile\\-Apps|APIs\\-Google|Crawl4AI|Crawlspace|FeedFetcher\\-Google|FirecrawlAgent|Google\\-CWS|Google\\-InspectionTool|Google\\-Safety|Google\\-Site\\-Verification|GoogleProducer|Lightpanda|Mediapartners\\-Google|Scrapy|Screaming\\ Frog\\ SEO\\ Spider|wpbot|YandexDirect|YandexMetrika|YandexScreenshotBot|YandexWebmaster)","dataset":"(AI2Bot|Ai2Bot\\-Dolma|aiHitBot|AwarioRssBot|AwarioSmartBot|CCBot|Diffbot|EchoboxBot|ImagesiftBot|img2dataset|LAIONDownloader|omgili|omgilibot|panscient\\.com|Thinkbot|VelenPublicWebCrawler|YaK)","preview":"(facebookexternalhit|GoogleMessages|Slackbot|Slackbot\\-LinkExpanding)","seo":"(AhrefsBot|AhrefsSiteAudit|barkrowler|DataForSeoBot|dotbot|MJ12bot|rogerbot|SemrushBot|SemrushBot\\-BA|SemrushBot\\-ESI|SemrushBot\\-FT|SemrushBot\\-SI|SemrushBot\\-SWA|SEOkicks|serpstatbot|SiteAuditBot|SplitSignalBot)","archive":"(archive\\.org_bot|ia_archiver)"}},"crawlers":[{"slug":"adsbot-google","name":"AdsBot-Google","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"AdsBot-Google","robots_token":"AdsBot-Google","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"own-token-only","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/adsbot-google.html","what_it_is":"Checks the quality of desktop landing pages for Google Ads. Google documents that it ignores the robots.txt * group with the ad publisher's permission, and obeys a group named for its own token.","cost_of_blocking":"Google Ads cannot score your landing pages, which lowers Ad Rank on the ads pointing at them. If you do not buy ads, blocking it costs nothing but bandwidth savings."},{"slug":"adsbot-google-mobile","name":"AdsBot-Google-Mobile","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"AdsBot-Google-Mobile","robots_token":"AdsBot-Google-Mobile","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"own-token-only","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/adsbot-google-mobile.html","what_it_is":"The mobile-web landing page checker for Google Ads. Same rules as AdsBot-Google: the * group does not apply to it, its own token does.","cost_of_blocking":"Mobile ad landing pages go unscored and the ads pointing at them rank worse. No effect on organic search."},{"slug":"adsbot-google-mobile-apps","name":"AdsBot-Google-Mobile-Apps","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"AdsBot-Google-Mobile-Apps","robots_token":"AdsBot-Google-Mobile-Apps","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"own-token-only","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/adsbot-google-mobile-apps.html","what_it_is":"Checks Android app landing pages for Google Ads. It obeys a group named for its own token and, per Google, follows the AdsBot-Google rules otherwise.","cost_of_blocking":"App-install ad landing pages go unscored. Nothing organic changes."},{"slug":"ahrefsbot","name":"AhrefsBot","operator":"Ahrefs","category":"seo","category_label":"SEO and backlink crawlers","ua":"AhrefsBot","robots_token":"AhrefsBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://ahrefs.com/robot","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/ahrefs-crawler.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ahrefsbot.html","what_it_is":"Ahrefs' backlink crawler, and one of the largest non-search crawlers on the web by request volume.","cost_of_blocking":"No user-facing effect. Ahrefs honours Crawl-delay, so rate-limiting is usually better than blocking."},{"slug":"ahrefssiteaudit","name":"AhrefsSiteAudit","operator":"Ahrefs","category":"seo","category_label":"SEO and backlink crawlers","ua":"AhrefsSiteAudit","robots_token":"AhrefsSiteAudit","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://ahrefs.com/robot","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/ahrefs-crawler.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ahrefssiteaudit.html","what_it_is":"Ahrefs' site-audit crawler, separate from AhrefsBot. Ahrefs documents that it obeys robots.txt by default, and that a verified site owner can ask for it to be allowed to ignore robots.txt on their own site so the audit can see disallowed sections.","cost_of_blocking":"Site owners auditing your domain get an incomplete report. Blocking it saves bandwidth and costs you nothing in search."},{"slug":"ai2bot","name":"AI2Bot","operator":"Allen Institute for AI","category":"dataset","category_label":"Corpus and dataset builders","ua":"AI2Bot","robots_token":"AI2Bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://allenai.org/crawler","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ai2bot.html","what_it_is":"The Allen Institute's crawler, gathering pages for open research corpora such as Dolma that underpin fully open models like OLMo.","cost_of_blocking":"Excluded from open research datasets. Worth a deliberate decision: this is the category where 'blocking AI' also blocks the open, auditable end of it."},{"slug":"ai2bot-dolma","name":"Ai2Bot-Dolma","operator":"Allen Institute for AI","category":"dataset","category_label":"Corpus and dataset builders","ua":"Ai2Bot-Dolma","robots_token":"Ai2Bot-Dolma","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://allenai.org/crawler","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ai2bot-dolma.html","what_it_is":"The variant of AI2's crawler named for the Dolma corpus specifically.","cost_of_blocking":"Same as AI2Bot: exclusion from an open, published training corpus."},{"slug":"aihitbot","name":"aiHitBot","operator":"aiHit","category":"dataset","category_label":"Corpus and dataset builders","ua":"aiHitBot","robots_token":"aiHitBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.aihitdata.com/about","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/aihitbot.html","what_it_is":"aiHit's automated collector, building a company dataset from public company websites.","cost_of_blocking":"Your company record in a B2B dataset goes stale. Documented as respecting robots.txt, so the rule works."},{"slug":"aiwebindex","name":"AIWebIndex","operator":"Lyrenth","category":"ai-search","category_label":"AI search crawlers","ua":"AIWebIndex","robots_token":"AIWebIndex","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://lyrenth.com/crawler-policy","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/aiwebindex.html","what_it_is":"Builds an index of public pages and serves them to AI agents as extracted readable text, with attribution and a link back. Lyrenth publishes a crawler policy stating it does not train foundation models on what it collects and that it obeys robots.txt.","cost_of_blocking":"Agents reading through this index stop seeing you — including the attribution and link back that make it a referral rather than a summary."},{"slug":"amazonbot","name":"Amazonbot","operator":"Amazon","category":"ai-search","category_label":"AI search crawlers","ua":"Amazonbot","robots_token":"Amazonbot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://developer.amazon.com/amazonbot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/amazonbot.html","what_it_is":"Amazon's crawler, feeding Alexa's ability to answer questions from the web and Amazon's own search and assistant products.","cost_of_blocking":"Alexa and Amazon's assistants stop answering from your pages. Verify with reverse DNS to crawl.amazonbot.amazon before trusting the user-agent."},{"slug":"andibot","name":"Andibot","operator":"Andi","category":"ai-search","category_label":"AI search crawlers","ua":"Andibot","robots_token":"Andibot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://andisearch.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/andibot.html","what_it_is":"The crawler for Andi, a small generative search assistant that summarises pages rather than listing them.","cost_of_blocking":"You disappear from another assistant's answers. Andi publishes no robots.txt statement."},{"slug":"anomura","name":"Anomura","operator":"Direqt","category":"ai-search","category_label":"AI search crawlers","ua":"Anomura","robots_token":"Anomura","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://direqt.ai","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/anomura.html","what_it_is":"Direqt's search crawler. It indexes the sites of Direqt's own publisher customers so their on-site chatbots can answer from them.","cost_of_blocking":"If you are the publisher, this breaks the assistant you put on your own pages. If you are not, it should not be crawling you."},{"slug":"anthropic-ai","name":"anthropic-ai","operator":"Anthropic","category":"ai-training","category_label":"AI training crawlers","ua":"anthropic-ai","robots_token":"anthropic-ai","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"n-a","docs":"https://support.anthropic.com/en/articles/8896518","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/anthropic-ai.html","what_it_is":"A legacy robots.txt token from before Anthropic consolidated on ClaudeBot. It is still widely present in robots.txt files and costs nothing to keep, but it is a control token rather than a bot you will see in logs.","cost_of_blocking":"None. Nothing crawls under this name today; keeping the rule is harmless insurance."},{"slug":"apis-google","name":"APIs-Google","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"APIs-Google","robots_token":"APIs-Google","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"own-token-only","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/apis-google.html","what_it_is":"Delivers push notifications for Google APIs to a webhook you registered. It is a special-case crawler: it ignores the robots.txt * group, because the fetch is a delivery to an address you asked it to deliver to.","cost_of_blocking":"Google API push notifications stop arriving at your endpoint. This only affects services you set up yourself; there is no search or AI consequence."},{"slug":"applebot","name":"Applebot","operator":"Apple","category":"search","category_label":"Search engines","ua":"Applebot","robots_token":"Applebot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://support.apple.com/en-us/119829","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/apple-applebot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/applebot.html","what_it_is":"Powers Siri, Spotlight and Safari suggestions. Blocking it is a search decision, not an AI decision — the AI decision has its own token.","cost_of_blocking":"You disappear from Siri, Spotlight and Safari search suggestions across Apple's install base."},{"slug":"applebot-extended","name":"Applebot-Extended","operator":"Apple","category":"ai-training","category_label":"AI training crawlers","ua":"(control token only — no crawler)","robots_token":"Applebot-Extended","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"n-a","docs":"https://support.apple.com/en-us/119829","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/applebot-extended.html","what_it_is":"Apple's counterpart to Google-Extended: a robots.txt token that withdraws consent for Apple Intelligence and Apple foundation-model training, without touching Applebot's search crawl.","cost_of_blocking":"Excluded from Apple Intelligence training. Siri, Spotlight and Safari suggestions are unaffected."},{"slug":"archive-org-bot","name":"archive.org_bot","operator":"Internet Archive","category":"archive","category_label":"Archivers","ua":"archive.org_bot","robots_token":"archive.org_bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://archive.org/details/archive.org_bot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/archive-org-bot.html","what_it_is":"The Wayback Machine's crawler. Preservation rather than AI, but it lands in the same 'is this bot welcome' decision and its output is a public corpus.","cost_of_blocking":"Your site stops being preserved. When it dies, it is gone. Consider this one separately from the AI question."},{"slug":"atlassian-bot","name":"atlassian-bot","operator":"Atlassian","category":"ai-search","category_label":"AI search crawlers","ua":"atlassian-bot","robots_token":"atlassian-bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://support.atlassian.com/organization-administration/docs/connect-custom-website-to-rovo/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/atlassian-bot.html","what_it_is":"Indexes a website so it can be searched and cited by Rovo, Atlassian's generative assistant inside Jira and Confluence. Atlassian's documentation walks a customer through editing robots.txt for it, which is as close to a compliance statement as this list gets.","cost_of_blocking":"Rovo cannot answer from your public documentation. If your customers live inside Atlassian tools, this is a support-deflection block."},{"slug":"awariorssbot","name":"AwarioRssBot","operator":"Awario","category":"dataset","category_label":"Corpus and dataset builders","ua":"AwarioRssBot","robots_token":"AwarioRssBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://awario.com/bots.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/awariorssbot.html","what_it_is":"The feed-reading half of Awario's pair, documented on the same page and under the same crawl-rate policy.","cost_of_blocking":"Your RSS updates stop reaching Awario's monitoring. Block both tokens or neither."},{"slug":"awariosmartbot","name":"AwarioSmartBot","operator":"Awario","category":"dataset","category_label":"Corpus and dataset builders","ua":"AwarioSmartBot","robots_token":"AwarioSmartBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://awario.com/bots.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/awariosmartbot.html","what_it_is":"Awario's brand-monitoring crawler. It documents one request per three seconds, honours Crawl-delay, and states it does not use consecutive IP blocks so identification is by user-agent only.","cost_of_blocking":"Mentions of brands on your pages stop being surfaced to the people monitoring them — including, quite possibly, your own."},{"slug":"baiduspider","name":"Baiduspider","operator":"Baidu","category":"search","category_label":"Search engines","ua":"Baiduspider","robots_token":"Baiduspider","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://help.baidu.com/question?prod_id=99&class=0&id=3001","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/baiduspider.html","what_it_is":"Baidu's search crawler, and the ingest path for Baidu's Ernie-backed answers.","cost_of_blocking":"Removal from Baidu Search, which matters only if you want Chinese-language traffic."},{"slug":"barkrowler","name":"Barkrowler","operator":"Babbar","category":"seo","category_label":"SEO and backlink crawlers","ua":"barkrowler","robots_token":"barkrowler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://babbar.tech/crawler","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/barkrowler.html","what_it_is":"Babbar's crawler, which builds the link graph behind their French-market SEO tooling.","cost_of_blocking":"You leave Babbar's index. No effect on search or assistants."},{"slug":"bedrockbot","name":"bedrockbot","operator":"Amazon","category":"ai-search","category_label":"AI search crawlers","ua":"bedrockbot","robots_token":"bedrockbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/webcrawl-data-source-connector.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/bedrockbot.html","what_it_is":"The web crawler an AWS customer points at URLs they chose, to build a knowledge base for a Bedrock application. AWS documents that it respects robots.txt and that the user-agent carries a per-customer suffix, so you can allow or refuse one customer's crawl by naming bedrockbot-UUID.","cost_of_blocking":"Companies building retrieval applications on Bedrock cannot include your pages. This is a RAG block, not a training block: nothing is being trained, but nothing can cite you either."},{"slug":"bingbot","name":"bingbot","operator":"Microsoft","category":"search","category_label":"Search engines","ua":"bingbot","robots_token":"bingbot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/bing-bingbot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/bingbot.html","what_it_is":"Bing's only crawler, and therefore also the crawler behind Microsoft Copilot's grounding. Microsoft's documented way to keep search indexing while refusing generative reuse is the nocache / noarchive robots meta directive, not a separate user-agent.","cost_of_blocking":"Very high and very wide: Bing, Copilot, DuckDuckGo and several assistants that resell Bing's index all lose you at once. Use nocache/noarchive rather than blocking."},{"slug":"bytespider","name":"Bytespider","operator":"ByteDance","category":"ai-training","category_label":"AI training crawlers","ua":"Bytespider","robots_token":"Bytespider","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"disputed","docs":"https://www.bytespider.net/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/bytespider.html","what_it_is":"ByteDance's crawler, associated with training data collection for Doubao and related models. Repeatedly reported by CDNs and site operators as the highest-volume AI crawler on the web and as inconsistent about robots.txt.","cost_of_blocking":"Little to lose. If you want it gone, expect to block by user-agent at the edge rather than to ask politely in robots.txt."},{"slug":"ccbot","name":"CCBot","operator":"Common Crawl","category":"dataset","category_label":"Corpus and dataset builders","ua":"CCBot","robots_token":"CCBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://commoncrawl.org/ccbot","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/commoncrawl-ccbot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ccbot.html","what_it_is":"Common Crawl's corpus builder. It trains nothing itself, but its archive is an input to most open and many closed LLM training sets, which makes it the highest-leverage single entry on this list.","cost_of_blocking":"Future Common Crawl snapshots exclude you, so downstream training sets lose you too — but only going forward. Existing snapshots are permanent and blocking today does not retract them."},{"slug":"chatgpt-agent","name":"ChatGPT Agent","operator":"OpenAI","category":"user-fetch","category_label":"User-triggered fetchers","ua":"ChatGPT Agent","robots_token":"ChatGPT-User","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://platform.openai.com/docs/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/openai-chatgpt-user.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/chatgpt-agent.html","what_it_is":"ChatGPT's agent mode driving a real browser: it navigates and interacts with sites to finish a multi-step task a user gave it. OpenAI governs it with the ChatGPT-User token and the ChatGPT-User prefix list rather than a token of its own, so the robots rule and the address check are the same ones.","cost_of_blocking":"Agentic tasks a user asked for — booking, comparing, filling a form on your site — fail. This is the fetch that ends in a transaction, so it is the most expensive user-triggered block on this list."},{"slug":"chatgpt-user","name":"ChatGPT-User","operator":"OpenAI","category":"user-fetch","category_label":"User-triggered fetchers","ua":"ChatGPT-User","robots_token":"ChatGPT-User","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://platform.openai.com/docs/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/openai-chatgpt-user.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/chatgpt-user.html","what_it_is":"Fetches a single page at the moment a user or a ChatGPT agent asks for it — a pasted link, a browsing step, an Operator task. One human intent, one request. OpenAI states these fetches are not used for training.","cost_of_blocking":"ChatGPT cannot open your pages when a user explicitly asks it to. The user sees a fetch failure. This is usually the last bot anyone means to block."},{"slug":"claude-searchbot","name":"Claude-SearchBot","operator":"Anthropic","category":"ai-search","category_label":"AI search crawlers","ua":"Claude-SearchBot","robots_token":"Claude-SearchBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://support.anthropic.com/en/articles/8896518","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/claude-searchbot.html","what_it_is":"Indexes pages so Claude's web search can find and cite them. Separate token from the training crawler, so search visibility and training consent are independent decisions.","cost_of_blocking":"You stop appearing in Claude's search results and citations."},{"slug":"claude-user","name":"Claude-User","operator":"Anthropic","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Claude-User","robots_token":"Claude-User","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://support.anthropic.com/en/articles/8896518","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/claude-user.html","what_it_is":"Fetches a page because a Claude user asked Claude to read it, at that moment.","cost_of_blocking":"Claude reports a fetch failure to a user who asked for your page by name."},{"slug":"claude-web","name":"Claude-Web","operator":"Anthropic","category":"ai-search","category_label":"AI search crawlers","ua":"Claude-Web","robots_token":"Claude-Web","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"n-a","docs":"https://support.anthropic.com/en/articles/8896518","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/claude-web.html","what_it_is":"An earlier Anthropic token for user-facing web access, superseded by Claude-User and Claude-SearchBot. Kept here because it appears in most published robots.txt templates.","cost_of_blocking":"None in practice. Retain the rule; expect no traffic."},{"slug":"claudebot","name":"ClaudeBot","operator":"Anthropic","category":"ai-training","category_label":"AI training crawlers","ua":"ClaudeBot","robots_token":"ClaudeBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://support.anthropic.com/en/articles/8896518","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/claudebot.html","what_it_is":"Anthropic's bulk crawler, gathering pages that may be used to train Claude models.","cost_of_blocking":"Content excluded from training data for future Claude models. No effect on Claude's ability to fetch a link a user gives it."},{"slug":"cloudflare-autorag","name":"Cloudflare-AutoRAG","operator":"Cloudflare","category":"ai-search","category_label":"AI search crawlers","ua":"Cloudflare-AutoRAG","robots_token":"Cloudflare-AutoRAG","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://developers.cloudflare.com/ai-search/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/cloudflare-autorag.html","what_it_is":"The crawler behind Cloudflare's AI Search / AutoRAG, which indexes a website into a retrieval index for an application. Cloudflare's own documentation warns that a bot-blocking rule on your zone will also stop this crawler and tells you to allow-list it.","cost_of_blocking":"Applications built on Cloudflare AI Search cannot retrieve your pages. If you are the one building the index over your own site, blocking it breaks your own product."},{"slug":"cohere-ai","name":"cohere-ai","operator":"Cohere","category":"user-fetch","category_label":"User-triggered fetchers","ua":"cohere-ai","robots_token":"cohere-ai","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://cohere.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/cohere-ai.html","what_it_is":"Cohere's fetcher, used when its assistant products need a page.","cost_of_blocking":"Cohere-powered assistants cannot read your pages on request."},{"slug":"cohere-training-data-crawler","name":"cohere-training-data-crawler","operator":"Cohere","category":"ai-training","category_label":"AI training crawlers","ua":"cohere-training-data-crawler","robots_token":"cohere-training-data-crawler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://cohere.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/cohere-training-data-crawler.html","what_it_is":"Cohere's separately-named bulk crawler for model training data, split out so consent for training and consent for retrieval can differ.","cost_of_blocking":"Excluded from Cohere model training."},{"slug":"cotoyogi","name":"Cotoyogi","operator":"ROIS-DS","category":"ai-training","category_label":"AI training crawlers","ua":"Cotoyogi","robots_token":"Cotoyogi","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://ds.rois.ac.jp/en_center8/en_crawler/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/cotoyogi.html","what_it_is":"A crawler run by ROIS-DS, a Japanese inter-university research organisation, collecting Japanese-language text for AI training. It publishes a crawler page in English and Japanese.","cost_of_blocking":"Your Japanese-language content is left out of an academic training corpus."},{"slug":"crawl4ai","name":"Crawl4AI","operator":"Crawl4AI project","category":"tool","category_label":"Tools and frameworks","ua":"Crawl4AI","robots_token":"Crawl4AI","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://github.com/unclecode/crawl4ai","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/crawl4ai.html","what_it_is":"An open-source LLM-oriented crawler and scraper library, run by whoever installs it. Like Scrapy, the default user-agent identifies the software and says nothing about who is behind the request.","cost_of_blocking":"You block a library, not an operator: the rule catches a researcher and a bulk scraper equally, and anyone who edits one config line is not caught at all."},{"slug":"crawlspace","name":"Crawlspace","operator":"Crawlspace","category":"tool","category_label":"Tools and frameworks","ua":"Crawlspace","robots_token":"Crawlspace","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://crawlspace.dev","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/crawlspace.html","what_it_is":"A crawling platform: customers run their own crawls on it to feed agents, RAG pipelines and structured-data workflows. Like Firecrawl, the party behind any given request is the customer, not the platform.","cost_of_blocking":"Whatever any Crawlspace customer was building over your pages stops working. Volume and intent vary per customer, so this is a rate-limit decision more than a consent one."},{"slug":"dataforseobot","name":"DataForSeoBot","operator":"DataForSEO","category":"seo","category_label":"SEO and backlink crawlers","ua":"DataForSeoBot","robots_token":"DataForSeoBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://dataforseo.com/dataforseo-bot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/dataforseobot.html","what_it_is":"Builds the backlink and SERP datasets DataForSEO resells through its API, so one crawl reaches many downstream tools.","cost_of_blocking":"You leave a dataset that a long tail of SEO products is built on. No user-facing effect."},{"slug":"diffbot","name":"Diffbot","operator":"Diffbot","category":"dataset","category_label":"Corpus and dataset builders","ua":"Diffbot","robots_token":"Diffbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://docs.diffbot.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/diffbot.html","what_it_is":"Extracts structured records from pages to build a commercial knowledge graph that is resold and used for retrieval and training.","cost_of_blocking":"Your facts stop entering a widely-licensed knowledge graph. Whether that is a loss depends on whether you want to be a machine-readable entity."},{"slug":"dotbot","name":"DotBot","operator":"Moz","category":"seo","category_label":"SEO and backlink crawlers","ua":"dotbot","robots_token":"dotbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://moz.com/help/moz-procedures/crawlers/dotbot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/dotbot.html","what_it_is":"Moz's crawler for Link Explorer. Moz documents that it respects robots.txt and that dotbot is the token to name.","cost_of_blocking":"You leave Moz's link index, so Domain Authority and link reports about your site get thinner. Nothing a reader or an assistant sees changes."},{"slug":"duckassistbot","name":"DuckAssistBot","operator":"DuckDuckGo","category":"ai-search","category_label":"AI search crawlers","ua":"DuckAssistBot","robots_token":"DuckAssistBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/duckassistbot.html","what_it_is":"Fetches pages so DuckAssist can generate and cite answers inside DuckDuckGo.","cost_of_blocking":"No DuckAssist answers or citations from your site. Ordinary DuckDuckGo results are unaffected."},{"slug":"duckduckbot","name":"DuckDuckBot","operator":"DuckDuckGo","category":"search","category_label":"Search engines","ua":"DuckDuckBot","robots_token":"DuckDuckBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://duckduckgo.com/duckduckgo-help-pages/results/duckduckbot/","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/duckduckgo-duckduckbot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/duckduckbot.html","what_it_is":"DuckDuckGo's own crawler. Note that the bulk of DuckDuckGo's web results come from Bing, so blocking bingbot removes you from DuckDuckGo whether or not you allow this one.","cost_of_blocking":"Limited on its own; the real DuckDuckGo lever is bingbot."},{"slug":"echoboxbot","name":"EchoboxBot","operator":"Echobox","category":"dataset","category_label":"Corpus and dataset builders","ua":"EchoboxBot","robots_token":"EchoboxBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://echobox.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/echoboxbot.html","what_it_is":"Collects data supporting Echobox's AI-driven social and email distribution products, which publishers use to schedule and target their own content.","cost_of_blocking":"Publishers using Echobox get worse scheduling decisions about your articles. No compliance statement is published."},{"slug":"exasearchbot","name":"ExaSearchBot","operator":"Exa","category":"ai-search","category_label":"AI search crawlers","ua":"ExaSearchBot","robots_token":"ExaSearchBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://exa.ai","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/exasearchbot.html","what_it_is":"Exa's crawler. It discovers and indexes public pages so they can be retrieved and cited through Exa's search API, which is one of the common retrieval backends behind agent frameworks.","cost_of_blocking":"Agents built on Exa's API stop finding you. Exa publishes no statement about robots.txt compliance, so treat the rule as a request."},{"slug":"facebookbot","name":"FacebookBot","operator":"Meta","category":"ai-training","category_label":"AI training crawlers","ua":"FacebookBot","robots_token":"FacebookBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/facebookbot.html","what_it_is":"Meta's older speech- and language-corpus crawler, largely superseded by meta-externalagent but still listed as a valid robots token.","cost_of_blocking":"Negligible today. Keep the rule; expect little traffic."},{"slug":"facebookexternalhit","name":"facebookexternalhit","operator":"Meta","category":"preview","category_label":"Link preview fetchers","ua":"facebookexternalhit","robots_token":"facebookexternalhit","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/facebookexternalhit.html","what_it_is":"The link unfurler: it reads your Open Graph tags when somebody shares your URL on a Meta property.","cost_of_blocking":"Severe and usually accidental. Your links share as bare grey boxes with no title, image or description across Facebook, Instagram, Messenger and WhatsApp. Almost nobody means to block this."},{"slug":"factset-spyderbot","name":"Factset_spyderbot","operator":"FactSet","category":"ai-training","category_label":"AI training crawlers","ua":"Factset_spyderbot","robots_token":"Factset_spyderbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.factset.com/ai","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/factset-spyderbot.html","what_it_is":"FactSet's crawler, collecting data used in AI model training for its financial data and analytics products.","cost_of_blocking":"Exclusion from a financial-data vendor's corpus. Relevant mostly to companies whose filings and disclosures are being read."},{"slug":"feedfetcher-google","name":"FeedFetcher-Google","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"FeedFetcher-Google","robots_token":"FeedFetcher-Google","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/feedfetcher-google.html","what_it_is":"Crawls RSS and Atom feeds for Google News and WebSub. It is a user-triggered fetcher, and Google documents that those generally ignore robots.txt because a person asked for the fetch. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"Feed-driven Google products stop seeing your updates. A robots.txt rule will not stop it — block by user-agent at the edge if you mean it."},{"slug":"firecrawlagent","name":"FirecrawlAgent","operator":"Firecrawl","category":"tool","category_label":"Tools and frameworks","ua":"FirecrawlAgent","robots_token":"FirecrawlAgent","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://docs.firecrawl.dev/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/firecrawlagent.html","what_it_is":"A hosted scrape-to-markdown service that LLM applications call to read pages. The requester is whoever is building on it, not Firecrawl itself, so volume and intent vary wildly.","cost_of_blocking":"Applications built on Firecrawl cannot read your pages. This is increasingly how agents fetch the web, so it is a bigger block than its name suggests."},{"slug":"google-agent","name":"Google-Agent","operator":"Google","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Google-Agent","robots_token":"Google-Agent","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered-agents.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-agent.html","what_it_is":"Agents hosted on Google infrastructure navigating the web and taking actions on a user's request. Google names one prefix list for it — user-triggered-agents.json — and is separately experimenting with Web Bot Auth under the identity https://agent.bot.goog.","cost_of_blocking":"Google-hosted agents cannot complete a task on your site for a user who asked them to. This is the agentic-commerce fetch: blocking it removes you from what an assistant can actually do rather than from what it can say."},{"slug":"google-cloudvertexbot","name":"Google-CloudVertexBot","operator":"Google","category":"ai-search","category_label":"AI search crawlers","ua":"Google-CloudVertexBot","robots_token":"Google-CloudVertexBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-cloudvertexbot.html","what_it_is":"Crawls a site on behalf of a Vertex AI Agent Builder customer who is building an agent over that site. It only visits sites the customer has asked it to.","cost_of_blocking":"Third parties can no longer build Vertex AI agents that read your site. Irrelevant to Google Search."},{"slug":"google-cws","name":"Google-CWS","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"Google-CWS","robots_token":"Google-CWS","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-cws.html","what_it_is":"The Chrome Web Store fetcher. It requests the URLs a developer put in the metadata of a Chrome extension or theme. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"Chrome Web Store listings that point at your pages cannot fetch them. Relevant only if you publish extensions."},{"slug":"google-extended","name":"Google-Extended","operator":"Google","category":"ai-training","category_label":"AI training crawlers","ua":"(control token only — no crawler)","robots_token":"Google-Extended","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"n-a","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-extended.html","what_it_is":"Not a crawler. A robots.txt token that tells Google whether pages Googlebot already fetched may be used to train and ground Gemini. You will never see it in an access log; disallowing it changes what Google does with content it fetched under a different name.","cost_of_blocking":"You are excluded from Gemini grounding and Gemini training. Google Search ranking and indexing are explicitly unaffected. This is the cleanest 'no training, keep my search traffic' lever that exists."},{"slug":"google-gemininotebook","name":"Google-GeminiNotebook","operator":"Google","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Google-GeminiNotebook","robots_token":"Google-GeminiNotebook","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-gemininotebook.html","what_it_is":"Fetches a URL a Gemini Notebook (formerly NotebookLM) user added as a source to their notebook. The former agent string Google-NotebookLM is documented as supported until August 2026. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"A user who deliberately added your page as a research source gets nothing. This is a citation-shaped fetch, not a training crawl."},{"slug":"google-inspectiontool","name":"Google-InspectionTool","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"Google-InspectionTool","robots_token":"Google-InspectionTool","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-inspectiontool.html","what_it_is":"The fetcher behind Search Console's URL Inspection and the Rich Results Test. It runs when a site owner clicks a button.","cost_of_blocking":"Your own Search Console live tests stop working. Blocking this only hurts you."},{"slug":"google-pinpoint","name":"Google-Pinpoint","operator":"Google","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Google-Pinpoint","robots_token":"Google-Pinpoint","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-pinpoint.html","what_it_is":"Fetches individual URLs that a Pinpoint user — usually a journalist or researcher — added as a source to their own document collection. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"A researcher who explicitly added your page to a collection cannot load it."},{"slug":"google-read-aloud","name":"Google-Read-Aloud","operator":"Google","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Google-Read-Aloud","robots_token":"Google-Read-Aloud","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-read-aloud.html","what_it_is":"Fetches a page so Google can read it out loud with text-to-speech, at the moment a user asks. Formerly google-speakr. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"A reader who asked Google to read your page aloud — often someone using it for accessibility — gets an error instead."},{"slug":"google-safety","name":"Google-Safety","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"Google-Safety","robots_token":"Google-Safety","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-safety.html","what_it_is":"Google's abuse-investigation fetcher: malware review, phishing reports and similar. Google documents that it ignores robots.txt entirely, and a robots.txt rule for it does nothing.","cost_of_blocking":"Nothing you can control. The rule is ignored by design; listing the token is documentation, not enforcement."},{"slug":"google-site-verification","name":"Google-Site-Verification","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"Google-Site-Verification","robots_token":"Google-Site-Verification","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/google-site-verification.html","what_it_is":"Fetches the token file or meta tag that proves you own a site, when you click verify in Search Console. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"Your own Search Console verification fails. Blocking this only ever hurts the person doing the blocking."},{"slug":"googlebot","name":"Googlebot","operator":"Google","category":"search","category_label":"Search engines","ua":"Googlebot","robots_token":"Googlebot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-googlebot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googlebot.html","what_it_is":"The classic search crawler. It is also the crawler behind AI Overviews: Google does not run a separate bot for them, which is why the only AI opt-out is the Google-Extended token and not a Googlebot block.","cost_of_blocking":"Total. You leave Google Search. Never block this to avoid AI use; use Google-Extended instead."},{"slug":"googlebot-image","name":"Googlebot-Image","operator":"Google","category":"search","category_label":"Search engines","ua":"Googlebot-Image","robots_token":"Googlebot-Image","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-googlebot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googlebot-image.html","what_it_is":"Image indexing for Google Images. A separate token so you can leave images out of search without leaving search.","cost_of_blocking":"Your images stop appearing in Google Images."},{"slug":"googlebot-news","name":"Googlebot-News","operator":"Google","category":"search","category_label":"Search engines","ua":"Googlebot-News","robots_token":"Googlebot-News","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-googlebot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googlebot-news.html","what_it_is":"A robots.txt token controlling inclusion in Google News. It does not have its own user-agent string; the fetch arrives as Googlebot.","cost_of_blocking":"Removal from Google News, with normal Search unaffected."},{"slug":"googlebot-video","name":"Googlebot-Video","operator":"Google","category":"search","category_label":"Search engines","ua":"Googlebot-Video","robots_token":"Googlebot-Video","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-googlebot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googlebot-video.html","what_it_is":"The video half of Googlebot. It crawls video files and the pages around them for Google Video search, and it is matched by a robots.txt group for Googlebot as well as by its own token.","cost_of_blocking":"Your videos leave Google video search. A rule for Googlebot already covers it, so blocking this token alone is usually a mistake of precision rather than of intent."},{"slug":"googlemessages","name":"GoogleMessages","operator":"Google","category":"preview","category_label":"Link preview fetchers","ua":"GoogleMessages","robots_token":"GoogleMessages","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googlemessages.html","what_it_is":"Generates the link preview when somebody sends one of your URLs in Google Messages. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"Your links appear as bare URLs with no title or image in Google Messages chats. A preview fetcher is almost never the one you meant to block."},{"slug":"googleother","name":"GoogleOther","operator":"Google","category":"ai-training","category_label":"AI training crawlers","ua":"GoogleOther","robots_token":"GoogleOther","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googleother.html","what_it_is":"A generic fetcher used by Google product teams for one-off crawls and research, including data collection that does not belong to Search.","cost_of_blocking":"No effect on Search indexing. Blocks internal Google research and product fetches."},{"slug":"googleother-image","name":"GoogleOther-Image","operator":"Google","category":"ai-training","category_label":"AI training crawlers","ua":"GoogleOther-Image","robots_token":"GoogleOther-Image","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googleother-image.html","what_it_is":"The image variant of GoogleOther: one-off fetches by Google product and research teams that are not Search. It also answers to a GoogleOther group in robots.txt.","cost_of_blocking":"Google teams outside Search stop fetching your images. Image Search itself is unaffected — that is Googlebot-Image."},{"slug":"googleother-video","name":"GoogleOther-Video","operator":"Google","category":"ai-training","category_label":"AI training crawlers","ua":"GoogleOther-Video","robots_token":"GoogleOther-Video","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googleother-video.html","what_it_is":"The video variant of GoogleOther, used for internal Google fetches that do not belong to Search.","cost_of_blocking":"No effect on Search or on Google Video search. Blocks internal Google research fetches of your video files."},{"slug":"googleproducer","name":"GoogleProducer","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"GoogleProducer","robots_token":"GoogleProducer","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-user-triggered.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/googleproducer.html","what_it_is":"Google Publisher Center: fetches the feeds a publisher explicitly supplied for Google News landing pages. Google publishes fetcher addresses in two files — user-triggered-fetchers.json and user-triggered-fetchers-google.json — and does not say per fetcher which one applies, so verification means checking both; this index mirrors both.","cost_of_blocking":"Your own Google News landing pages stop updating. Only publishers who configured Publisher Center are affected."},{"slug":"gptbot","name":"GPTBot","operator":"OpenAI","category":"ai-training","category_label":"AI training crawlers","ua":"GPTBot","robots_token":"GPTBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://platform.openai.com/docs/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/openai-gptbot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/gptbot.html","what_it_is":"OpenAI's bulk crawler. Pages it fetches may be used to train future OpenAI foundation models. It is not the bot that puts you in ChatGPT's search results, and blocking it does not remove you from them.","cost_of_blocking":"Your content is excluded from training data for future OpenAI models. No effect on ChatGPT search visibility, on citations, or on links a user pastes into ChatGPT."},{"slug":"ia-archiver","name":"ia_archiver","operator":"Internet Archive","category":"archive","category_label":"Archivers","ua":"ia_archiver","robots_token":"ia_archiver","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://archive.org/details/archive.org_bot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/ia-archiver.html","what_it_is":"The legacy Alexa/Internet Archive token, still present in most robots.txt files and still occasionally honoured.","cost_of_blocking":"Negligible today; retain for tidiness."},{"slug":"icc-crawler","name":"ICC-Crawler","operator":"NICT","category":"ai-training","category_label":"AI training crawlers","ua":"ICC-Crawler","robots_token":"ICC-Crawler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.nict.go.jp/en/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/icc-crawler.html","what_it_is":"Operated by NICT, Japan's national information and communications research institute. The collected data supports AI research and, per the operator, is also provided to third parties including commercial companies.","cost_of_blocking":"You are excluded from a national research corpus and from the commercial redistributions of it. This is a dataset-shaped block: one refusal, many downstream effects."},{"slug":"imagesiftbot","name":"ImagesiftBot","operator":"Hive AI","category":"dataset","category_label":"Corpus and dataset builders","ua":"ImagesiftBot","robots_token":"ImagesiftBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://imagesift.com/about","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/imagesiftbot.html","what_it_is":"Crawls images for Hive AI's reverse-image and dataset products. Image-heavy sites see this one long before they see the text crawlers.","cost_of_blocking":"Your images stop entering an image dataset and reverse-image index."},{"slug":"img2dataset","name":"img2dataset","operator":"LAION / img2dataset","category":"dataset","category_label":"Corpus and dataset builders","ua":"img2dataset","robots_token":"img2dataset","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://github.com/rom1504/img2dataset","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/img2dataset.html","what_it_is":"The tool used to turn image-URL lists such as LAION's into downloaded training sets. It is run by whoever is building a dataset, not by a single operator.","cost_of_blocking":"Your images are skipped when someone materialises an image-text dataset that references them."},{"slug":"isscyberriskcrawler","name":"ISSCyberRiskCrawler","operator":"ISS Corporate Solutions","category":"ai-training","category_label":"AI training crawlers","ua":"ISSCyberRiskCrawler","robots_token":"ISSCyberRiskCrawler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"disputed","docs":"https://iss-cyber.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/isscyberriskcrawler.html","what_it_is":"Crawls in order to train models that score a company's cyber risk. The ai.robots.txt dataset records the operator as not respecting robots.txt; ISS publishes no compliance statement of its own.","cost_of_blocking":"A rule here is a statement of intent. Your organisation's public footprint still gets scored — by a model trained on everybody else."},{"slug":"kagibot","name":"Kagibot","operator":"Kagi","category":"search","category_label":"Search engines","ua":"Kagibot","robots_token":"Kagibot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://kagi.com/bot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/kagibot.html","what_it_is":"The crawler for Kagi, a paid, ad-free search engine with its own index and its own assistant.","cost_of_blocking":"You leave Kagi's index. Kagi's users are paying to search and skew technical; per visitor this is an expensive block."},{"slug":"klaviyoaibot","name":"KlaviyoAIBot","operator":"Klaviyo","category":"ai-search","category_label":"AI search crawlers","ua":"KlaviyoAIBot","robots_token":"KlaviyoAIBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://help.klaviyo.com/hc/en-us/articles/40496146232219","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/klaviyoaibot.html","what_it_is":"Fetches pages from domains a Klaviyo customer has explicitly connected to their own account, to power Klaviyo's Kai customer agent. It is scoped to connected domains rather than the open web.","cost_of_blocking":"If the connected domain is yours, blocking this breaks the agent you configured. If it is not, this bot should not be reaching you at all."},{"slug":"laiondownloader","name":"LAIONDownloader","operator":"LAION / img2dataset","category":"dataset","category_label":"Corpus and dataset builders","ua":"LAIONDownloader","robots_token":"LAIONDownloader","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"by-design-no","docs":"https://laion.ai/faq/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/laiondownloader.html","what_it_is":"LAION's downloader, used to materialise the image and text datasets the non-profit publishes for machine-learning research. LAION's own FAQ is the source for its robots.txt position.","cost_of_blocking":"Your media is skipped when an open research dataset is built from URL lists. Once a dataset is published, a later block does not remove you from it."},{"slug":"lightpanda","name":"Lightpanda","operator":"Lightpanda","category":"tool","category_label":"Tools and frameworks","ua":"Lightpanda","robots_token":"Lightpanda","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://lightpanda.io/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/lightpanda.html","what_it_is":"A purpose-built headless browser for AI and automation — a runtime, not an operator. Whether robots.txt is honoured is left to whoever runs it, which is what its maintainers say themselves.","cost_of_blocking":"You block a browser, not a company: the same rule stops a scraper and a legitimate automation a customer of yours is running."},{"slug":"linguee-bot","name":"Linguee Bot","operator":"Linguee","category":"ai-training","category_label":"AI training crawlers","ua":"Linguee Bot","robots_token":"Linguee Bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"disputed","docs":"https://www.linguee.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/linguee-bot.html","what_it_is":"Gathers bilingual text for Linguee's translation corpus and the machine translation trained on it. Recorded in the ai.robots.txt dataset as not respecting robots.txt.","cost_of_blocking":"Multilingual pages stop feeding a translation corpus. If your site is translated, being in it is usually a benefit."},{"slug":"mediapartners-google","name":"Mediapartners-Google","operator":"Google","category":"tool","category_label":"Tools and frameworks","ua":"Mediapartners-Google","robots_token":"Mediapartners-Google","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"own-token-only","docs":"https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/mediapartners-google.html","what_it_is":"The AdSense crawler. It reads a page so AdSense can choose relevant ads for it, and it is a special-case crawler that ignores the robots.txt * group.","cost_of_blocking":"Pages it cannot read get generic, lower-value AdSense ads or none at all. This is the one block on this list that costs you money directly if you run AdSense."},{"slug":"meta-externalagent","name":"meta-externalagent","operator":"Meta","category":"ai-training","category_label":"AI training crawlers","ua":"meta-externalagent","robots_token":"meta-externalagent","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/meta-externalagent.html","what_it_is":"Meta's AI crawler, gathering training data for Llama and Meta AI. It replaced the older FacebookBot name for this purpose.","cost_of_blocking":"Excluded from Meta AI training. Link previews on Facebook, Instagram and WhatsApp are unaffected — those are a different bot."},{"slug":"meta-externalfetcher","name":"meta-externalfetcher","operator":"Meta","category":"user-fetch","category_label":"User-triggered fetchers","ua":"meta-externalfetcher","robots_token":"meta-externalfetcher","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/meta-externalfetcher.html","what_it_is":"Fetches a page when a Meta AI user asks about a specific link.","cost_of_blocking":"Meta AI cannot read pages users hand it."},{"slug":"meta-webindexer","name":"Meta-WebIndexer","operator":"Meta","category":"ai-search","category_label":"AI search crawlers","ua":"Meta-WebIndexer","robots_token":"Meta-WebIndexer","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/meta-webindexer.html","what_it_is":"Per Meta's crawler documentation, Meta-WebIndexer navigates the web to improve the quality of Meta AI's search results. It is a third Meta token alongside Meta-ExternalAgent and Meta-ExternalFetcher, and the newest of them.","cost_of_blocking":"You leave the index Meta AI answers from across Facebook, Instagram and WhatsApp — the largest assistant install base there is. A robots.txt that names the two older Meta tokens does not cover this one."},{"slug":"mistralai-user","name":"MistralAI-User","operator":"Mistral AI","category":"user-fetch","category_label":"User-triggered fetchers","ua":"MistralAI-User","robots_token":"MistralAI-User","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://docs.mistral.ai/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/mistralai-user.html","what_it_is":"Fetches a page when a Le Chat user asks Mistral's assistant to read it.","cost_of_blocking":"Le Chat cannot open links your readers give it."},{"slug":"mj12bot","name":"MJ12bot","operator":"Majestic","category":"seo","category_label":"SEO and backlink crawlers","ua":"MJ12bot","robots_token":"MJ12bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://mj12bot.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/mj12bot.html","what_it_is":"Majestic's link-graph crawler, run as a distributed community project. Majestic states plainly that it cannot restrict the bot to a fixed set of addresses, and offers a pre-arranged ident string in the request headers instead.","cost_of_blocking":"You leave the Majestic backlink index. No search or AI effect. It supports Crawl-delay, which is usually the better answer than a block."},{"slug":"mojeekbot","name":"MojeekBot","operator":"Mojeek","category":"search","category_label":"Search engines","ua":"MojeekBot","robots_token":"MojeekBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.mojeek.com/bot.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/mojeekbot.html","what_it_is":"Mojeek's crawler. Mojeek runs one of the few genuinely independent web indexes — not a front end over Bing or Google — so it is one of the few blocks that removes you from an index nobody else can put you back into. Its documentation states it obeys the first record whose User-Agent contains MojeekBot, falling back to *.","cost_of_blocking":"You leave an independent index that other privacy-focused search products draw on. Small traffic, disproportionate long-term cost to web plurality."},{"slug":"oai-searchbot","name":"OAI-SearchBot","operator":"OpenAI","category":"ai-search","category_label":"AI search crawlers","ua":"OAI-SearchBot","robots_token":"OAI-SearchBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://platform.openai.com/docs/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/openai-searchbot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/oai-searchbot.html","what_it_is":"Builds the index ChatGPT search answers from. Content it collects is used for retrieval and citation, not for model training.","cost_of_blocking":"High. Blocking this removes you from ChatGPT search results and from the source links ChatGPT shows. This is the single most expensive block on this list for anyone who wants to be cited by an assistant."},{"slug":"omgili","name":"omgili","operator":"Webz.io","category":"dataset","category_label":"Corpus and dataset builders","ua":"omgili","robots_token":"omgili","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://webz.io/blog/machine-learning/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/omgili.html","what_it_is":"The older robots token for the same Webz.io collection, still honoured and still worth listing.","cost_of_blocking":"Same as omgilibot."},{"slug":"omgilibot","name":"omgilibot","operator":"Webz.io","category":"dataset","category_label":"Corpus and dataset builders","ua":"omgilibot","robots_token":"omgilibot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://webz.io/blog/machine-learning/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/omgilibot.html","what_it_is":"Webz.io's crawler, collecting web and forum text sold as datasets, including to model builders.","cost_of_blocking":"Exclusion from a commercial dataset resold to third parties."},{"slug":"panscient","name":"Panscient","operator":"Panscient","category":"dataset","category_label":"Corpus and dataset builders","ua":"panscient.com","robots_token":"panscient.com","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://panscient.com/faq.htm","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/panscient.html","what_it_is":"Compiles structured data about businesses and business professionals using machine learning. Panscient's FAQ states it obeys robots.txt.","cost_of_blocking":"Your company pages stop feeding a business-data product. No effect on search or assistants."},{"slug":"perplexity-user","name":"Perplexity-User","operator":"Perplexity","category":"user-fetch","category_label":"User-triggered fetchers","ua":"Perplexity-User","robots_token":"Perplexity-User","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"by-design-no","docs":"https://docs.perplexity.ai/guides/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/perplexity-user.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/perplexity-user.html","what_it_is":"Fetches a page because a Perplexity user asked for it. Perplexity documents that this fetch is user-initiated and is therefore not governed by robots.txt — a robots rule will not stop it, by stated policy.","cost_of_blocking":"Not controllable via robots.txt. If you must stop it, verify by the published IP ranges and block at the edge — and accept that users who ask for your page get an error."},{"slug":"perplexitybot","name":"PerplexityBot","operator":"Perplexity","category":"ai-search","category_label":"AI search crawlers","ua":"PerplexityBot","robots_token":"PerplexityBot","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://docs.perplexity.ai/guides/bots","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/perplexity-bot.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/perplexitybot.html","what_it_is":"Builds Perplexity's search index. Perplexity is citation-heavy by product design, so inclusion here converts to referral traffic more directly than most AI surfaces.","cost_of_blocking":"You stop being indexed and cited by Perplexity, and lose the referral clicks its citations produce."},{"slug":"petalbot","name":"PetalBot","operator":"Huawei","category":"search","category_label":"Search engines","ua":"PetalBot","robots_token":"PetalBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://aspiegel.com/petalbot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/petalbot.html","what_it_is":"Huawei's crawler for Petal Search, shipped as the default search on Huawei devices.","cost_of_blocking":"Removal from Petal Search. Frequently blocked for volume rather than for policy."},{"slug":"phindbot","name":"PhindBot","operator":"Phind","category":"ai-search","category_label":"AI search crawlers","ua":"PhindBot","robots_token":"PhindBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.phind.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/phindbot.html","what_it_is":"Phind is an answer engine for developers that combines live web search with its own models. This is the crawler behind those answers.","cost_of_blocking":"You stop being cited in answers to technical questions — which, for documentation and reference sites, is the exact audience most worth keeping."},{"slug":"pinterestbot","name":"Pinterestbot","operator":"Pinterest","category":"search","category_label":"Search engines","ua":"Pinterestbot","robots_token":"Pinterestbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://help.pinterest.com/en/business/article/pinterest-crawler","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/pinterestbot.html","what_it_is":"Pinterest's crawler. It indexes pages so people can find them on Pinterest and re-reads product pages to keep price and title on a Pin current. Pinterest states that content it crawls is not used to train their Canvas image generation model.","cost_of_blocking":"Pins pointing at your site go stale — wrong prices, dead links — and new content stops being indexed. For a retailer this is one of the more expensive blocks on the list."},{"slug":"poseidon-research-crawler","name":"Poseidon Research Crawler","operator":"Poseidon Research","category":"ai-training","category_label":"AI training crawlers","ua":"Poseidon Research Crawler","robots_token":"Poseidon Research Crawler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.poseidonresearch.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/poseidon-research-crawler.html","what_it_is":"A crawler run by Poseidon Research, a lab working on interpretability research for AI systems.","cost_of_blocking":"Exclusion from an interpretability research corpus. No published compliance statement."},{"slug":"qualifiedbot","name":"QualifiedBot","operator":"Qualified","category":"ai-search","category_label":"AI search crawlers","ua":"QualifiedBot","robots_token":"QualifiedBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.qualified.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/qualifiedbot.html","what_it_is":"Analyses a customer's website so Qualified's AI sales chatbots can answer questions about it in context.","cost_of_blocking":"A chatbot on a site that licensed the product loses context. If that site is yours, this block is self-inflicted."},{"slug":"quillbot","name":"QuillBot","operator":"QuillBot","category":"ai-training","category_label":"AI training crawlers","ua":"QuillBot","robots_token":"QuillBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://quillbot.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/quillbot.html","what_it_is":"Operated by QuillBot as part of its writing, paraphrasing and AI-detection products. The dataset also records a second token, quillbot.com, for the same operator.","cost_of_blocking":"Exclusion from QuillBot's corpus. No compliance statement is published, so the rule is a request."},{"slug":"qwantbot","name":"Qwantbot","operator":"Qwant","category":"search","category_label":"Search engines","ua":"Qwantbot","robots_token":"Qwantbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://help.qwant.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/qwantbot.html","what_it_is":"Qwant's crawler. Qwant documents that the string Qwantbot always appears in its user-agents whatever the crawler version, which is what makes a substring match safe here.","cost_of_blocking":"You leave the index behind Qwant, the French privacy-focused engine, and the products that federate it."},{"slug":"qwantbot-news","name":"Qwantbot-news","operator":"Qwant","category":"search","category_label":"Search engines","ua":"Qwantbot-news","robots_token":"Qwantbot-news","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://help.qwant.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/qwantbot-news.html","what_it_is":"The news variant of Qwant's crawler, documented alongside the main one and carrying the same Qwantbot substring.","cost_of_blocking":"Your articles stop appearing in Qwant News. A rule for Qwantbot as a substring already catches both."},{"slug":"reflectionbot","name":"Reflectionbot","operator":"Reflection AI","category":"ai-training","category_label":"AI training crawlers","ua":"Reflectionbot","robots_token":"Reflectionbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://reflection.ai/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/reflectionbot.html","what_it_is":"An undocumented crawler whose user-agent links to Reflection AI, a company building AI models. The link in the user-agent is the only public statement of purpose that exists.","cost_of_blocking":"Unknown by construction — which is itself the reason some people block it. Nothing user-facing depends on it."},{"slug":"rogerbot","name":"rogerbot","operator":"Moz","category":"seo","category_label":"SEO and backlink crawlers","ua":"rogerbot","robots_token":"rogerbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://moz.com/help/moz-procedures/crawlers/rogerbot","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/rogerbot.html","what_it_is":"Moz's Campaign crawler, which audits a site its own owner registered. Moz states there is no IP range for it — identification is by user-agent only.","cost_of_blocking":"Moz Pro site audits of your domain stop. If the domain is yours and you use Moz, blocking this breaks your own reports."},{"slug":"sbintuitionsbot","name":"SBIntuitionsBot","operator":"SB Intuitions","category":"ai-training","category_label":"AI training crawlers","ua":"SBIntuitionsBot","robots_token":"SBIntuitionsBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.sbintuitions.co.jp/en/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/sbintuitionsbot.html","what_it_is":"SB Intuitions is SoftBank's Japanese LLM lab; this crawler gathers data used in that model development and in information analysis. The operator publishes a dedicated bot page.","cost_of_blocking":"Your content is excluded from a Japanese-language foundation-model corpus. Nothing user-facing changes."},{"slug":"scrapy","name":"Scrapy","operator":"Scrapy project","category":"tool","category_label":"Tools and frameworks","ua":"Scrapy","robots_token":"Scrapy","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://scrapy.org/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/scrapy.html","what_it_is":"Not an operator: the default user-agent of the most common Python crawling framework. Anyone can be behind it. Modern Scrapy obeys robots.txt by default, which is why the default UA is still worth a rule.","cost_of_blocking":"You block a very large tail of unattributed one-off crawlers, and also every well-behaved researcher who did not change the default."},{"slug":"screaming-frog-seo-spider","name":"Screaming Frog SEO Spider","operator":"Screaming Frog","category":"tool","category_label":"Tools and frameworks","ua":"Screaming Frog SEO Spider","robots_token":"Screaming Frog SEO Spider","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.screamingfrog.co.uk/seo-spider/user-agent/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/screaming-frog-seo-spider.html","what_it_is":"Not an operator: desktop crawling software that anybody can point at any site. The default user-agent identifies the tool, not who is running it, and the operator of the moment is whoever pressed start.","cost_of_blocking":"You block a consultant auditing your own site as often as you block a stranger. Treat it as a rate-limit question, not a consent one — and note that the user-agent is configurable, so a block is advisory."},{"slug":"semrushbot","name":"SemrushBot","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot","robots_token":"SemrushBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot.html","what_it_is":"Semrush's backlink and keyword crawler. It is not an AI crawler, but it is usually in the top three by volume on any site, and it is the cheapest block on this list.","cost_of_blocking":"Your competitors' Semrush reports get thinner, and so do yours. No user-facing effect."},{"slug":"semrushbot-ba","name":"SemrushBot-BA","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot-BA","robots_token":"SemrushBot-BA","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-ba.html","what_it_is":"The Backlink Audit crawler. It re-checks links pointing at a customer's site, which means it lands on the sites doing the linking.","cost_of_blocking":"None to you. It costs the site being audited a little accuracy in their backlink report."},{"slug":"semrushbot-esi","name":"SemrushBot-ESI","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot-ESI","robots_token":"SemrushBot-ESI","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-esi.html","what_it_is":"The crawler for Semrush Enterprise Site Intelligence, the enterprise tier's own site analysis.","cost_of_blocking":"Enterprise customers lose analysis of your domain. Nothing user-facing."},{"slug":"semrushbot-ft","name":"SemrushBot-FT","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot-FT","robots_token":"SemrushBot-FT","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-ft.html","what_it_is":"Fetches full text for the Plagiarism Checker and similar text-comparison tools.","cost_of_blocking":"Your text stops being compared against other people's submissions — which also means copies of your text are less likely to be caught."},{"slug":"semrushbot-ocob","name":"SemrushBot-OCOB","operator":"Semrush","category":"ai-training","category_label":"AI training crawlers","ua":"SemrushBot-OCOB","robots_token":"SemrushBot-OCOB","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-ocob.html","what_it_is":"Semrush's separately-tokenised crawler for its AI content tooling, split out so SEO crawling and AI reuse can be answered differently.","cost_of_blocking":"Exclusion from Semrush's AI corpus, with its SEO crawl unaffected."},{"slug":"semrushbot-si","name":"SemrushBot-SI","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot-SI","robots_token":"SemrushBot-SI","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-si.html","what_it_is":"Fetches pages for the On Page SEO Checker and similar advisory tools.","cost_of_blocking":"Nothing user-facing. Semrush customers lose on-page suggestions for pages on your domain."},{"slug":"semrushbot-swa","name":"SemrushBot-SWA","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SemrushBot-SWA","robots_token":"SemrushBot-SWA","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/semrushbot-swa.html","what_it_is":"Checks whether a URL is reachable, for the SEO Writing Assistant. One request per URL a writer references, not a crawl.","cost_of_blocking":"Writers using Semrush's assistant see your links reported as unreachable."},{"slug":"seokicks","name":"SEOkicks","operator":"SEOkicks","category":"seo","category_label":"SEO and backlink crawlers","ua":"SEOkicks","robots_token":"SEOkicks","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.seokicks.de/robot.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/seokicks.html","what_it_is":"A German backlink index. Its documentation names SEOkicks as the user-agent to use in robots.txt.","cost_of_blocking":"You leave a regional backlink index. Nothing else changes."},{"slug":"serpstatbot","name":"serpstatbot","operator":"Serpstat","category":"seo","category_label":"SEO and backlink crawlers","ua":"serpstatbot","robots_token":"serpstatbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://serpstatbot.com/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/serpstatbot.html","what_it_is":"Serpstat's backlink crawler. It documents support for Crawl-delay up to 20 seconds, including a delay set on the * group.","cost_of_blocking":"You leave Serpstat's link index. Try Crawl-delay first — they honour it, and a slow crawler is cheaper to keep than to fight."},{"slug":"seznambot","name":"SeznamBot","operator":"Seznam","category":"search","category_label":"Search engines","ua":"SeznamBot","robots_token":"SeznamBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://napoveda.seznam.cz/en/seznamzbozi/subject-matter-crawler/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/seznambot.html","what_it_is":"Seznam's crawler — the dominant search engine in the Czech Republic and one of the few national engines with its own index.","cost_of_blocking":"Removal from Seznam. Also removes you from its IndexNow endpoint's usefulness."},{"slug":"shapbot","name":"ShapBot","operator":"Parallel","category":"ai-search","category_label":"AI search crawlers","ua":"ShapBot","robots_token":"ShapBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://docs.parallel.ai/features/crawler","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/shapbot.html","what_it_is":"Parallel's crawler. It collects and structures web content to power the search, extraction and deep-research APIs that Parallel sells to agent builders.","cost_of_blocking":"Agents using Parallel's research API lose you as a source. Parallel documents robots.txt compliance, so a rule works."},{"slug":"sidetrade-indexer-bot","name":"Sidetrade indexer bot","operator":"Sidetrade","category":"ai-training","category_label":"AI training crawlers","ua":"Sidetrade indexer bot","robots_token":"Sidetrade indexer bot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.sidetrade.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/sidetrade-indexer-bot.html","what_it_is":"Sidetrade extracts web data for a range of uses including training its AI products for order-to-cash and customer-data work.","cost_of_blocking":"Exclusion from a commercial B2B dataset. The operator publishes no robots.txt statement, so treat the rule as a request rather than a control."},{"slug":"siteauditbot","name":"SiteAuditBot","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SiteAuditBot","robots_token":"SiteAuditBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/siteauditbot.html","what_it_is":"Semrush's Site Audit crawler: it walks a site a customer owns and reports technical SEO problems. Semrush names it as the token to block for that product.","cost_of_blocking":"Site Audit reports on your own domain stop working. If a customer is auditing your site with your permission, blocking this breaks their tooling and nothing of yours."},{"slug":"slackbot","name":"Slackbot","operator":"Slack","category":"preview","category_label":"Link preview fetchers","ua":"Slackbot","robots_token":"Slackbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://api.slack.com/robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/slackbot.html","what_it_is":"The other half of Slack's pair: the agent that reads robots.txt and handles Slack's non-unfurl fetches. Slack documents both strings on one page.","cost_of_blocking":"Slack stops being able to read your robots.txt, which is a strange thing to want. Block the link expander instead if that is the goal."},{"slug":"slackbot-linkexpanding","name":"Slackbot-LinkExpanding","operator":"Slack","category":"preview","category_label":"Link preview fetchers","ua":"Slackbot-LinkExpanding","robots_token":"Slackbot-LinkExpanding","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://api.slack.com/robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/slackbot-linkexpanding.html","what_it_is":"Fetches a page to build the unfurl card when somebody pastes your link into Slack. One paste, one fetch.","cost_of_blocking":"Your links appear in Slack as bare URLs. Inside working teams that quietly costs you clicks, and nothing is gained."},{"slug":"splitsignalbot","name":"SplitSignalBot","operator":"Semrush","category":"seo","category_label":"SEO and backlink crawlers","ua":"SplitSignalBot","robots_token":"SplitSignalBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://www.semrush.com/bot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/splitsignalbot.html","what_it_is":"Runs SEO A/B tests on a customer's own site with the SplitSignal tool.","cost_of_blocking":"A site owner's own A/B testing stops. Only relevant on domains whose owner uses the product."},{"slug":"storebot-google","name":"Storebot-Google","operator":"Google","category":"search","category_label":"Search engines","ua":"Storebot-Google","robots_token":"Storebot-Google","verification_method":"published-ranges","verification_label":"published IP ranges","respects_robots_txt":"documented","docs":"https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers","ip_ranges":"https://www.pathwren.workers.dev/c/rubygems-registry/ip-ranges/google-special.json","url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/storebot-google.html","what_it_is":"Checks shopping and checkout flows for Google's shopping surfaces.","cost_of_blocking":"Product listings may lose shopping-specific enrichment. Irrelevant to non-commerce sites."},{"slug":"terracotta","name":"TerraCotta","operator":"Ceramic AI","category":"ai-search","category_label":"AI search crawlers","ua":"TerraCotta","robots_token":"TerraCotta","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://ceramic.ai/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/terracotta.html","what_it_is":"Ceramic AI's crawler, which indexes public content for a web-scale search API aimed at LLMs and agents.","cost_of_blocking":"You are absent from another agent-facing retrieval index. Ceramic documents that it obeys robots.txt."},{"slug":"thinkbot","name":"Thinkbot","operator":"Thinkbot","category":"dataset","category_label":"Corpus and dataset builders","ua":"Thinkbot","robots_token":"Thinkbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"disputed","docs":"https://www.thinkbot.agency","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/thinkbot.html","what_it_is":"Collects pages for analysis of how sites are adopting AI and automation. The ai.robots.txt dataset records the operator as not respecting robots.txt.","cost_of_blocking":"Exclusion from a market-research dataset. Expect to enforce this at the edge rather than in robots.txt."},{"slug":"tiktokspider","name":"TikTokSpider","operator":"ByteDance","category":"ai-training","category_label":"AI training crawlers","ua":"TikTokSpider","robots_token":"TikTokSpider","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"disputed","docs":"https://www.bytespider.net/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/tiktokspider.html","what_it_is":"A second ByteDance crawler identifying with TikTok, collecting page content for the same family of models.","cost_of_blocking":"Little to lose unless TikTok search referral matters to you."},{"slug":"timpibot","name":"Timpibot","operator":"Timpi","category":"search","category_label":"Search engines","ua":"Timpibot","robots_token":"Timpibot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://timpi.io/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/timpibot.html","what_it_is":"A distributed crawler building an independent search index outside the Google/Bing duopoly.","cost_of_blocking":"Absence from a small independent index."},{"slug":"velenpublicwebcrawler","name":"VelenPublicWebCrawler","operator":"Hunter (Velen)","category":"dataset","category_label":"Corpus and dataset builders","ua":"VelenPublicWebCrawler","robots_token":"VelenPublicWebCrawler","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://velen.io/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/velenpublicwebcrawler.html","what_it_is":"Hunter's crawler, written in Go, building business datasets and machine-learning models from public pages. Its page states it follows robots.txt and meta directives and never fetches more than one page every two seconds.","cost_of_blocking":"Your company pages stop feeding a B2B contact and company dataset. The crawl rate it documents makes this one of the cheapest visitors to simply allow."},{"slug":"webzio-extended","name":"Webzio-Extended","operator":"Webz.io","category":"ai-training","category_label":"AI training crawlers","ua":"Webzio-Extended","robots_token":"Webzio-Extended","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://webz.io/blog/machine-learning/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/webzio-extended.html","what_it_is":"Webz.io's opt-out token specifically for AI training reuse, in the pattern Google and Apple established.","cost_of_blocking":"Your content is excluded from the AI-training tier of Webz.io's product while ordinary collection continues."},{"slug":"wpbot","name":"wpbot","operator":"QuantumCloud","category":"tool","category_label":"Tools and frameworks","ua":"wpbot","robots_token":"wpbot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.quantumcloud.com","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/wpbot.html","what_it_is":"Supports the AI Chatbot for WordPress plugin: it reads pages so the plugin can answer from a site's own content. The operator provides an opt-out through a form rather than through robots.txt.","cost_of_blocking":"A WordPress site running that plugin loses its own content as an answer source. Only relevant where the plugin is installed."},{"slug":"yak","name":"YaK","operator":"Meltwater","category":"dataset","category_label":"Corpus and dataset builders","ua":"YaK","robots_token":"YaK","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"undocumented","docs":"https://www.meltwater.com/en/suite/consumer-intelligence","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yak.html","what_it_is":"Meltwater's crawler, feeding the live data stream behind its media-monitoring and consumer-intelligence suite.","cost_of_blocking":"Your content stops appearing in Meltwater's media monitoring — which is how PR teams find out you were mentioned. Some publishers want to be in it."},{"slug":"yandexadditional","name":"YandexAdditional","operator":"Yandex","category":"ai-training","category_label":"AI training crawlers","ua":"YandexAdditional","robots_token":"YandexAdditional","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexadditional.html","what_it_is":"The token that controls whether already-indexed pages may appear in Search with Yandex AI answers. Yandex's table says it makes no indexing requests of its own — it exists so a site can opt out of the generative answer without leaving the index.","cost_of_blocking":"You disappear from Yandex's AI answers while staying in Yandex Search. This is Yandex's equivalent of Google-Extended, and it is the cheap opt-out most people are looking for."},{"slug":"yandexadditionalbot","name":"YandexAdditionalBot","operator":"Yandex","category":"ai-training","category_label":"AI training crawlers","ua":"YandexAdditionalBot","robots_token":"YandexAdditionalBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexadditionalbot.html","what_it_is":"The second token Yandex publishes for the same AI-answers opt-out. Both names appear in Yandex's own robot list, so a robots.txt that names only one of them is half a policy.","cost_of_blocking":"Same as YandexAdditional: out of Yandex's AI answers, still in Yandex Search. Name both tokens or neither."},{"slug":"yandexblogs","name":"YandexBlogs","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexBlogs","robots_token":"YandexBlogs","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexblogs.html","what_it_is":"Yandex's blog-search robot; it indexes post comments as well as posts.","cost_of_blocking":"Comment threads and blog posts stop being findable through Yandex blog search."},{"slug":"yandexbot","name":"YandexBot","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexBot","robots_token":"YandexBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/robot-workings/check-yandex-robots.html","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexbot.html","what_it_is":"Yandex's search crawler, which also feeds Alice and Yandex's generative answers.","cost_of_blocking":"Removal from Yandex Search. Verify with reverse DNS to a yandex.ru, yandex.net or yandex.com host — YandexBot is among the most-spoofed user-agents there is."},{"slug":"yandexcalendar","name":"YandexCalendar","operator":"Yandex","category":"user-fetch","category_label":"User-triggered fetchers","ua":"YandexCalendar","robots_token":"YandexCalendar","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexcalendar.html","what_it_is":"Downloads calendar files a user subscribed to. Yandex notes these files are often in directories that are disallowed for indexing, which is why the general rules are not applied.","cost_of_blocking":"Users who subscribed to a calendar you publish stop receiving updates."},{"slug":"yandexcombot","name":"YandexComBot","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexComBot","robots_token":"YandexComBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexcombot.html","what_it_is":"Indexes content for Yandex search in languages other than Russian. Yandex documents that it can index content when there is no explicit robot-specific restriction — a * group is not one.","cost_of_blocking":"You leave Yandex's non-Russian index. A rule naming this token is the only one that works on it."},{"slug":"yandexdirect","name":"YandexDirect","operator":"Yandex","category":"tool","category_label":"Tools and frameworks","ua":"YandexDirect","robots_token":"YandexDirect","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexdirect.html","what_it_is":"Reads the content of Yandex Advertising Network partner pages to work out their topic so relevant ads can be matched. Documented as not taking the general robots.txt rules into account.","cost_of_blocking":"Ads on your pages become less relevant and earn less. Relevant only if you monetise with Yandex's network."},{"slug":"yandexfavicons","name":"YandexFavicons","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexFavicons","robots_token":"YandexFavicons","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexfavicons.html","what_it_is":"Downloads your favicon so Yandex can show it beside your result. Documented as not taking the general robots.txt rules into account.","cost_of_blocking":"Your results in Yandex lose their icon. Cosmetic, and a * rule will not achieve it anyway."},{"slug":"yandeximages","name":"YandexImages","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexImages","robots_token":"YandexImages","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandeximages.html","what_it_is":"Indexes images for Yandex Images. Yandex's robot table marks it as taking the general robots.txt rules into account.","cost_of_blocking":"Your images leave Yandex Images, which is a large share of image search in Russian-speaking markets."},{"slug":"yandexmarket","name":"YandexMarket","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexMarket","robots_token":"YandexMarket","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexmarket.html","what_it_is":"The robot behind Yandex Market, Yandex's shopping comparison service. Version 1.0 is documented as obeying the general rules; version 2.0 is documented as not.","cost_of_blocking":"Your products stop being listed and priced in Yandex Market. For a retailer in that market this is a revenue block, not a bandwidth one."},{"slug":"yandexmedia","name":"YandexMedia","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexMedia","robots_token":"YandexMedia","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexmedia.html","what_it_is":"Indexes multimedia data for Yandex. Takes the general robots.txt rules into account.","cost_of_blocking":"Your multimedia content stops appearing in Yandex's media surfaces."},{"slug":"yandexmetrika","name":"YandexMetrika","operator":"Yandex","category":"tool","category_label":"Tools and frameworks","ua":"YandexMetrika","robots_token":"YandexMetrika","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"by-design-no","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexmetrika.html","what_it_is":"Yandex Metrica's own fetcher. Two of its versions — the 2.0 yabs01 availability checker and the 4.0 CSS cache for Webvisor — are documented in Yandex's table as not using robots.txt at all.","cost_of_blocking":"Nothing you can enforce through robots.txt. If you run Metrica, its session replay loses your stylesheets and renders your pages wrong."},{"slug":"yandexmobilebot","name":"YandexMobileBot","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexMobileBot","robots_token":"YandexMobileBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexmobilebot.html","what_it_is":"Decides whether a page's layout is suitable for mobile devices. Yandex's table marks it as NOT taking the general robots.txt rules into account, so a * group does not stop it — a group named YandexMobileBot does.","cost_of_blocking":"Yandex loses its mobile-friendliness signal for your pages, which affects how they are ranked and rendered on phones."},{"slug":"yandexrenderresourcesbot","name":"YandexRenderResourcesBot","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexRenderResourcesBot","robots_token":"YandexRenderResourcesBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexrenderresourcesbot.html","what_it_is":"Loads the CSS, JavaScript and images Yandex needs to render a page. Yandex documents the exact rule: it ignores robots.txt for a resource when the HTML page using it is allowed, and does not fetch the resource when that page is disallowed.","cost_of_blocking":"Yandex renders your pages without their stylesheets or scripts and ranks what it sees. This is the classic accidental self-inflicted ranking loss."},{"slug":"yandexscreenshotbot","name":"YandexScreenshotBot","operator":"Yandex","category":"tool","category_label":"Tools and frameworks","ua":"YandexScreenshotBot","robots_token":"YandexScreenshotBot","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"own-token-only","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexscreenshotbot.html","what_it_is":"Takes a screenshot of a page. Documented as not taking the general robots.txt rules into account.","cost_of_blocking":"Yandex surfaces that show a page thumbnail show nothing for you."},{"slug":"yandexvideo","name":"YandexVideo","operator":"Yandex","category":"search","category_label":"Search engines","ua":"YandexVideo","robots_token":"YandexVideo","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexvideo.html","what_it_is":"Indexes video for Yandex video search. Obeys the general robots.txt rules per Yandex's own table.","cost_of_blocking":"Removal from Yandex video search. Note that a second robot, YandexVideoParser, does the same job and is documented as NOT taking the general rules into account."},{"slug":"yandexwebmaster","name":"YandexWebmaster","operator":"Yandex","category":"tool","category_label":"Tools and frameworks","ua":"YandexWebmaster","robots_token":"YandexWebmaster","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yandexwebmaster.html","what_it_is":"The fetcher behind Yandex Webmaster, the console a site owner uses to inspect their own site.","cost_of_blocking":"Your own Yandex Webmaster checks stop working. Blocking this only hurts you."},{"slug":"yeti","name":"Yeti","operator":"Naver","category":"search","category_label":"Search engines","ua":"Yeti","robots_token":"Yeti","verification_method":"reverse-dns","verification_label":"reverse DNS","respects_robots_txt":"documented","docs":"https://searchadvisor.naver.com/guide/seo-basic-crawl","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/yeti.html","what_it_is":"Naver's crawler. Naver is South Korea's largest search portal and runs its own index and its own generative answers.","cost_of_blocking":"Removal from Naver, which is most of Korean search."},{"slug":"youbot","name":"YouBot","operator":"You.com","category":"ai-search","category_label":"AI search crawlers","ua":"YouBot","robots_token":"YouBot","verification_method":"none","verification_label":"no published verification method","respects_robots_txt":"documented","docs":"https://about.you.com/youbot/","ip_ranges":null,"url":"https://www.pathwren.workers.dev/c/rubygems-registry/crawler/youbot.html","what_it_is":"You.com's crawler, feeding its AI search product and its search API.","cost_of_blocking":"Removal from You.com's index and from answers built on its API."}]}
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: ai-crawler-index
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 1.
|
|
4
|
+
version: 1.3.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Pathwren
|
|
@@ -32,10 +32,11 @@ licenses:
|
|
|
32
32
|
metadata:
|
|
33
33
|
homepage_uri: https://www.pathwren.workers.dev/c/rubygems-registry/
|
|
34
34
|
documentation_uri: https://www.pathwren.workers.dev/c/rubygems-registry/
|
|
35
|
+
source_code_uri: https://www.pathwren.workers.dev/c/rubygems-registry/git/ai-crawler-index.git/
|
|
35
36
|
changelog_uri: https://www.pathwren.workers.dev/c/rubygems-registry/changelog.html
|
|
36
37
|
data_source_uri: https://www.pathwren.workers.dev/c/rubygems-registry/data/agents.json
|
|
37
38
|
data_license: CC0-1.0
|
|
38
|
-
generated_at: '2026-09-
|
|
39
|
+
generated_at: '2026-09-05T17:25:28+00:00'
|
|
39
40
|
crawlers: '150'
|
|
40
41
|
rdoc_options: []
|
|
41
42
|
require_paths:
|