← KapaCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Kapa
Snapshot Oct 7, 2026 · 18:03 UTC · version 0.1.0
Collection source: downloaded plugin package.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Set up a Kapa web crawl source so a documentation site is ingested. Use when the user wants Kapa to answer from a website, documentation site or sitemap.",
"included_files": [],
"name": "kapa-setup-web-crawl",
"skill_md_contents": "---\nname: kapa-setup-web-crawl\ndescription: Set up a Kapa web crawl source so a documentation site is ingested. Use when the user wants Kapa to answer from a website, documentation site or sitemap.\n---\n\n# Set up a web crawl\n\nThree things happen in order: find the right pages, extract the right text from\nthem, then ingest. The first two are free and repeatable, the third spends the\nteam's quota. So judge the first two before doing the third.\n\nA source cannot be ingested until it has been previewed with the exact\nconfiguration you intend to ingest with. Changing the crawl config retires the\npreview, so a change means previewing again.\n\n## 1. Create the source\n\n`create_web_crawl_source` with `project` and `name`. Name it after the site, such as\n\"Acme docs\". Keep the returned source id.\n\n## 2. Say what to crawl\n\n`set_crawl_config` with `source_scrape` (the source id) and `urls_start`, the\npages the crawl begins from.\n\n- `urls_include` and `urls_exclude` narrow it. Use them when the site holds\n content the user does not want answered from, such as a blog or a changelog.\n- `enable_sitemap` follows the site's sitemap. It describes one site, so it only\n applies with a single start URL.\n- `render_js` is for a site that builds its content in the browser. It is\n slower, so leave it off until a preview shows thin pages.\n\nStart URLs must be public `http://` or `https://` addresses. A private or\nloopback host is rejected, since the crawler cannot reach it.\n\n## 3. Preview the crawl\n\n`preview_crawl` with the source `id` and `page_limit: 50`. Nothing is ingested\nand no quota is spent.\n\nThen `get_crawl_status` until it leaves `PENDING` and `IN_PROGRESS`. `SUCCESS`\nmeans it finished, `FAILURE` carries the reason. Poll every couple of seconds\nrather than in a tight loop.\n\n`list_preview_pages` with `source_scrape` to see what it found. **Read this\nlist and judge it**, since nothing else will:\n\n- Pages the user would not want answered from mean `urls_exclude` is needed.\n- A section that should be there and is not means `urls_start` or\n `urls_include` is wrong, or the site needs `render_js`.\n\nFix the config and preview again until the set of URLs is right. Only then move\non. A capped preview reports as much, so raise or drop the limit if 50 pages\nwere not representative.\n\nTo change the config, use `update_crawl_config`, not `set_crawl_config`.\n`set_crawl_config` creates one and refuses a second call. `update_crawl_config`\ntakes the **config id**, which `get_crawl_config` returns, not the source id.\n\n## 4. Settle the content selector\n\nThe selector picks the element holding the article text. Everything outside it,\nnavigation, sidebars, footers, is dropped. Getting this wrong degrades every\nanswer the source ever gives, and nothing errors.\n\n1. `detect_content_selector` with the source `id`. It recognises common\n documentation platforms and answers with settings, or null if it cannot tell.\n Call it once, not in a loop: it is rate limited. Use its `selector`,\n `selectors_exclude` and `classes_exclude`, and ignore `detected_by` and\n `platform`, which are metadata and must never be saved.\n2. `inspect_content_selector` with the `id` and a candidate `selector` renders\n it against real preview pages without saving. Start broad, with `main` or\n `article`, then narrow.\n3. **Read the extracted text, both its start and its end.** Navigation or\n \"related articles\" in the output means the selector is too broad. A nearly\n empty page means it is too narrow, or the page needs `render_js`. Repeat\n until only the article remains.\n4. `set_content_selector` with `source_scrape` and the selector you settled on.\n To change it later use `update_content_selector`, which takes the content\n config id rather than the source id.\n\nKeep the headings. Kapa chunks on the breadcrumb that headings form, so a\nselector that strips `h1` or `h2` quietly makes retrieval worse.\n\n`selectors_exclude` takes full CSS selectors. Use it for things inside the\narticle that are not article text, such as an edit link or a banner.\n\n## 5. Ingest\n\n`start_crawl` reads every page it finds and spends the team's quota, so confirm\nwith the user first.\n\nThis is the step that populates the project. A source that is configured but\nnever crawled answers nothing, so do not stop at step 4.\n\nAfterwards `get_crawl_status` follows it, and `cancel_crawl` stops it.\n\n## Changing a source that is already live\n\nEditing the configuration of a deployed source changes what it serves. Say so\nand get the user's agreement before saving.\n\nExtraction changes only take effect by ingesting again, so a new selector on a\nlive source does nothing until `start_crawl` re-runs.\n\n## When things go wrong\n\n**\"Crawl in progress\" or a 409.** A crawl is running, or approved pages are\nstill processing. Nothing clears it from here. Tell the user to wait, and do not\nretry in a loop.\n\n**No pages found.** Not an error, a result. The crawl config matched nothing:\ncheck the start URLs are reachable and that the include patterns are not too\nstrict.\n\n**Pages come back thin.** Their content is built in the browser. API reference\npages often are. Either turn on `render_js` and preview again, or exclude those\npaths and add an OpenAPI source instead, with `create_openapi_source` and\n`set_openapi_config`. A spec is structured, so it gives cleaner coverage than\nany crawl of the rendered page.\n\n**The preview will not start.** The error carries the reason. A failed preview\nleaves a draft source behind, so reuse it rather than creating another.\n\n## Shared Kapa workflow rules\n\nTools act as the connected user with that user's project permissions. Resolve the intended project and use only authorized data. Do not invent credentials, source IDs, filters, or tool results. Check the available tool schema before passing arguments.\n\nExplain and obtain approval for ingestion and its quota cost before saving a configuration that starts ingestion or calling `start_crawl`; existing explicit approval for that exact action is sufficient. Ask the user to choose source scope and filters. Validate credentials and discover accessible content before saving. Keep credentials out of visible results, logs, and exported artifacts. Use a secure credential input if the host provides one.\n\nFor a web source, preview the exact configuration and inspect the extracted article content before ingestion. Report queued, running, failed, and completed states accurately. If uncertain about Kapa behavior, use `search_kapa_docs` when available.\n"
}SHA-256 of public snapshot: 731d2149bdf2e773c6938e0640760e9a42ea7b9b2d36a6386cd27992a2545f52