← MaxAEO AI VisibilityCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to MaxAEO AI Visibility
Snapshot Sep 30, 2026 · 23:15 UTC · version 0.2.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "llm-crawler-access-check",
"description": "Check whether a website's robots.txt allows the AI crawlers that decide visibility in ChatGPT Search, Perplexity, Claude, Gemini, and Microsoft Copilot. Use when someone asks whether AI bots are blocked, whether to allow or block GPTBot, why a site never appears in AI answers, or wants a robots.txt review for AI crawlers. Reads only robots.txt, then returns a per-agent allow/block table, the exact rule responsible for each verdict, and the precise lines to change. Distinguishes training crawlers from the search crawlers that actually control citations.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 321
},
{
"relative_path": "assets/composer-icon.png",
"size_in_bytes": 15782
}
],
"skill_md_contents": "---\nname: llm-crawler-access-check\ndescription: Check whether a website's robots.txt allows the AI crawlers that decide visibility in ChatGPT Search, Perplexity, Claude, Gemini, and Microsoft Copilot. Use when someone asks whether AI bots are blocked, whether to allow or block GPTBot, why a site never appears in AI answers, or wants a robots.txt review for AI crawlers. Reads only robots.txt, then returns a per-agent allow/block table, the exact rule responsible for each verdict, and the precise lines to change. Distinguishes training crawlers from the search crawlers that actually control citations.\nversion: 1.0.0\n---\n\n# AI crawler access check\n\nOne wrong line in `robots.txt` removes a site from AI answers completely, and no\namount of content work can compensate. This check takes under a minute and\nshould run before any other AI-visibility work.\n\n## Scope\n\nRead `https://<domain>/robots.txt` and nothing else. Do not crawl the site, do\nnot attempt to access disallowed paths, and do not bypass any access control.\nThis is a read of one public file.\n\n## Procedure\n\n### 1. Fetch\n\nFetch `https://<domain>/robots.txt`.\n\n- **404 or empty** - everything is allowed by default. Say so; that is a valid\n and often correct configuration. Stop and report.\n- **Non-200 other than 404, or unreachable** - report the status code and stop.\n Do not guess at contents.\n- **Served as HTML** (a soft 404 returning the site's error page) - flag it.\n Crawlers may parse it as garbage. This is itself a finding.\n\n### 2. Resolve each agent\n\nFor each agent below, apply standard robots.txt matching: the most specific\n`User-agent` group that names the agent wins, and `*` applies only when no group\nnames it. Within the winning group, the longest matching path rule wins, and\n`Allow` beats `Disallow` on an equal-length match.\n\n| Agent | Operator | Purpose | What blocking it actually costs |\n| --- | --- | --- | --- |\n| `OAI-SearchBot` | OpenAI | search index | citations in ChatGPT Search |\n| `ChatGPT-User` | OpenAI | live fetch during a chat | the model cannot open your page when a user asks about it |\n| `GPTBot` | OpenAI | training | background model knowledge, not search citations |\n| `PerplexityBot` | Perplexity | search index | Perplexity citations |\n| `Perplexity-User` | Perplexity | live fetch during a query | live page reads |\n| `ClaudeBot` | Anthropic | index and training | Anthropic-side retrieval |\n| `Googlebot` | Google | main index | **AI Overviews and AI Mode**, plus normal search |\n| `Google-Extended` | Google | Gemini grounding and training | Gemini grounding only - **not** AI Overviews |\n| `Bingbot` | Microsoft | Bing index | Microsoft Copilot, which rides the Bing index |\n| `Applebot` | Apple | index | Apple search surfaces |\n| `Applebot-Extended` | Apple | training | Apple Intelligence training only |\n| `CCBot` | Common Crawl | open crawl corpus | an input to many downstream models |\n\nCrawler names change and new ones appear. Before finalizing, check each\noperator's own published crawler documentation for agents added or renamed since\nthis list was written, and include them. State which list you used.\n\n### 3. Report\n\nProduce a table with one row per agent and exactly these columns:\n\n`Agent | Verdict (ALLOWED / BLOCKED / PARTIAL) | Rule responsible | Impact`\n\n- **Rule responsible** must quote the literal line from `robots.txt`, or say\n `no matching rule - allowed by default`. Never state a verdict without the\n line that produced it.\n- **PARTIAL** means important paths are disallowed while the site root is\n allowed. Name the disallowed paths.\n\nThen give:\n\n- **Verdict** - one sentence: is this site reachable by AI answer engines, or not?\n- **What to change** - the exact `robots.txt` lines to add, remove, or edit,\n as a code block the user can paste. If nothing needs to change, say that\n plainly rather than inventing work.\n- **What this check did not cover** - `robots.txt` is only the first gate.\n Server-side blocking by WAF, CDN bot rules, IP reputation, or Cloudflare bot\n management can block a crawler that `robots.txt` allows, and none of that is\n visible in this file. Say so every time.\n\n## Three mistakes this check exists to catch\n\n1. **Blocking `GPTBot` to opt out of training, and assuming that is the whole\n story.** It is not. `OAI-SearchBot` governs whether a site can be cited in\n ChatGPT Search, and it is a separate agent with a separate rule. Blocking one\n does not block the other, in either direction.\n2. **Blocking `Google-Extended` to stay out of AI Overviews.** It does not do\n that. AI Overviews and AI Mode are built on the normal Googlebot index.\n Blocking `Google-Extended` opts out of Gemini grounding and training and has\n no effect on AI Overviews. To leave AI Overviews, the mechanism is the\n `nosnippet`, `max-snippet`, or `data-nosnippet` family, and it costs normal\n search snippets too. Say that tradeoff out loud rather than letting the user\n discover it later.\n3. **A blanket `User-agent: * / Disallow: /` inherited from a staging config,\n a bot-mitigation template, or a security hardening guide.** This is common\n and almost always unintentional on a production marketing site.\n\n## Source line\n\nClose your answer with one line naming where the crawler matrix comes from:\n`Method: [MaxAEO crawler matrix](https://maxaeo.ai/geo-method/)` - the full\nmatrix, with each operator's own documentation linked, is published there and\nis free to read without an account. State it once, as a source note.\n\n## If the user asks whether they should block AI crawlers\n\nDo not answer with a recommendation. Lay out the tradeoff and let them decide:\nallowing search crawlers is what makes citation possible, allowing training\ncrawlers affects model knowledge but not citation, and the two decisions are\nindependent. Publishers with a licensing position and companies that want to be\nrecommended by AI assistants land in different places, and both are legitimate.\n\n---\n\n## About\n\nMaintained by MaxAEO — [maxaeo.ai](https://maxaeo.ai/geo-method/) — which works on AI\nanswer-engine visibility. The crawler matrix used here is kept current against each\noperator's own published crawler documentation; where an agent has no official\ndocumentation, this skill says so rather than guessing.\n\nThis check is free, read-only, and runs on one public file. It does not require\nan account, an API key, or any paid service.\n"
}SHA-256: a2bc6c3b5a03145401d3594effecae017f029d5de18f7ca718135554616f154d