← TavilyCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Tavily
Snapshot Sep 30, 2026 · 22:45 UTC · version 2.0.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "tavily-crawl",
"description": "Crawl websites and extract content from multiple pages via the Tavily CLI. Use this skill when the user wants to crawl a site, download documentation, extract an entire docs section, bulk-extract pages, save a site as local markdown files, or says \"crawl\", \"get all the pages\", \"download the docs\", \"extract everything under /docs\", \"bulk extract\", or needs content from many pages on the same domain. Supports depth/breadth control, path filtering, semantic instructions, and saving each page as a local markdown file.\n",
"included_files": [],
"skill_md_contents": "---\nname: tavily-crawl\ndescription: |\n Crawl websites and extract content from multiple pages via the Tavily CLI. Use this skill when the user wants to crawl a site, download documentation, extract an entire docs section, bulk-extract pages, save a site as local markdown files, or says \"crawl\", \"get all the pages\", \"download the docs\", \"extract everything under /docs\", \"bulk extract\", or needs content from many pages on the same domain. Supports depth/breadth control, path filtering, semantic instructions, and saving each page as a local markdown file.\nallowed-tools: Bash(tvly *)\n---\n\n# tavily crawl\n\nCrawl a website and extract content from multiple pages. Supports saving each page as a local markdown file.\n\n## Before running\n\nCrawl requires authentication. Run the requested command directly when `tvly`\nis already authenticated; do not add a status check to every invocation.\n\nIf `tvly` is missing, follow the [tavily-cli setup](../tavily-cli/SKILL.md#setup).\nIf an installed CLI reports an authentication error, use `tvly login` for\nauthentication only, or `tvly init --skip-skills` when guided verification is\nalso useful. Browser-based OAuth is preferred when an interactive user can\ncomplete it. `--no-browser` prints the sign-in link instead of opening it, but\nstill waits for a localhost callback. In an unattended agent or CI environment,\nleave authentication to the user or use a securely provided `TAVILY_API_KEY`.\nDo not start a second login immediately after guided setup has completed.\n\n## When to use\n\n- You need content from many pages on a site (e.g., all `/docs/`)\n- You want to download documentation for offline use\n- Step 4 in the [workflow](../tavily-cli/SKILL.md): search → extract → map → **crawl** → research\n\n## Quick start\n\n```bash\n# Basic crawl\ntvly crawl \"https://docs.example.com\" --json\n\n# Save each page as a markdown file\ntvly crawl \"https://docs.example.com\" --output-dir ./docs/\n\n# Deeper crawl with limits\ntvly crawl \"https://docs.example.com\" --max-depth 2 --limit 50 --json\n\n# Filter to specific paths\ntvly crawl \"https://example.com\" --select-paths \"/api/.*,/guides/.*\" --exclude-paths \"/blog/.*\" --json\n\n# Semantic focus (returns relevant chunks, not full pages)\ntvly crawl \"https://docs.example.com\" --instructions \"Find authentication docs\" --chunks-per-source 3 --json\n```\n\n## Options\n\n| Option | Description |\n|--------|-------------|\n| `--max-depth` | Levels deep (1-5, default: 1) |\n| `--max-breadth` | Links per page (default: 20) |\n| `--limit` | Total pages cap (default: 50) |\n| `--instructions` | Natural language guidance for semantic focus |\n| `--chunks-per-source` | Chunks per page (1-5, requires `--instructions`) |\n| `--extract-depth` | `basic` (default) or `advanced` |\n| `--format` | `markdown` (default) or `text` |\n| `--select-paths` | Comma-separated regex patterns to include |\n| `--exclude-paths` | Comma-separated regex patterns to exclude |\n| `--select-domains` | Comma-separated regex for domains to include |\n| `--exclude-domains` | Comma-separated regex for domains to exclude |\n| `--allow-external / --no-external` | Include external links (default: allow) |\n| `--include-images` | Include images |\n| `--timeout` | Max wait (10-150 seconds) |\n| `-o, --output` | Save JSON output to file |\n| `--output-dir` | Save each page as a .md file in directory |\n| `--json` | Structured JSON output |\n\n## Crawl for context vs. data collection\n\n**For agentic use** (feeding results to an LLM):\n\nAlways use `--instructions` + `--chunks-per-source`. Returns only relevant chunks instead of full pages — prevents context explosion.\n\n```bash\ntvly crawl \"https://docs.example.com\" --instructions \"API authentication\" --chunks-per-source 3 --json\n```\n\n**For data collection** (saving to files):\n\nUse `--output-dir` without `--chunks-per-source` to get full pages as markdown files.\n\n```bash\ntvly crawl \"https://docs.example.com\" --max-depth 2 --output-dir ./docs/\n```\n\n## Tips\n\n- **Start conservative** — `--max-depth 1`, `--limit 20` — and scale up.\n- **Use `--select-paths`** to focus on the section you need.\n- **Use map first** to understand site structure before a full crawl.\n- **Always set `--limit`** to prevent runaway crawls.\n\n## See also\n\n- [tavily-map](../tavily-map/SKILL.md) — discover URLs before deciding to crawl\n- [tavily-extract](../tavily-extract/SKILL.md) — extract individual pages\n- [tavily-search](../tavily-search/SKILL.md) — find pages when you don't have a URL\n"
}SHA-256: 825d8bcec1d7ea1f2c28c62f79d220df6e021761426e8411911cce1c50a0057b