Public pages only
Homepage, robots.txt, sitemap.xml, llms.txt – no forms, no logins, no protected areas.
When a website blocks automated access, snaff identifies itself with its own, unique crawler. Allow it – and snaff analyzes protected sites in full.
Mozilla/5.0 (compatible; snaff-GEO-Checker/1.0; +https://snaff.ai/bot)
snaff is not a search or training crawler. It only fetches what a GEO analysis needs.
Homepage, robots.txt, sitemap.xml, llms.txt – no forms, no logins, no protected areas.
It only runs when a check is started or monitoring is set up – no mass crawling.
The user agent always contains “snaff-GEO-Checker” and links to this page – ideal for a targeted allow rule.
One WAF rule lets the snaff crawler past your bot protection – every other rule stays active.
Security → WAF → Custom rules → Create rule.
Field User Agent, operator contains, value snaff-GEO-Checker – or paste the expression below.
Tick “All remaining custom rules”, “Rate limiting rules”, “Managed rules” and – if available – “Super Bot Fight Mode”. Move the rule to the top.
(http.user_agent contains "snaff-GEO-Checker")
If your bot protection blocks snaff, it usually blocks ChatGPT, Claude and Perplexity too. Then your site can hardly be cited there.
Cloudflare: Security → Bots – turn off “Block AI Scrapers and Crawlers” or allow the crawlers you want under AI Crawl Control. Akamai, Imperva, DataDome, Sucuri: set them to “allow” in the bot manager.
GPTBotOAI-SearchBotChatGPT-UserClaudeBotClaude-SearchBotPerplexityBotGoogle-ExtendedApplebot-Extendedrobots.txt must be reachable without a browser check – otherwise crawlers can’t tell what they may do. snaff checks this automatically for every blocked site.
Questions or problems with the crawler: hello@snaff.ai
Check a website arrow_forward