Skills
Skill 3 of 8
Crawl websites and extract content from multiple pages via the Tavily CLI.
2 minutes · 340 words · 7 sections
Install
npx skills add tavily-ai/skills --skill tavily-crawlnpx skills add tavily-ai/skills/plugin marketplace add tavily-ai/skillsThe first command installs just this skill, by the name in its SKILL.md; the second installs the whole repository.
Crawl a website and extract content from multiple pages. Supports saving each page as a local markdown file.
Crawl requires authentication. Run the requested command directly when tvly
is already authenticated; do not add a status check to every invocation.
If tvly is missing, follow the tavily-cli setup.
If an installed CLI reports an authentication error, use tvly login for
authentication only, or tvly init --skip-skills when guided verification is
also useful. Browser-based OAuth is preferred when an interactive user can
complete it. --no-browser prints the sign-in link instead of opening it, but
still waits for a localhost callback. In an unattended agent or CI environment,
leave authentication to the user or use a securely provided TAVILY_API_KEY.
Do not start a second login immediately after guided setup has completed.
/docs/)# Basic crawl
tvly crawl "https://docs.example.com" --json
# Save each page as a markdown file
tvly crawl "https://docs.example.com" --output-dir ./docs/
# Deeper crawl with limits
tvly crawl "https://docs.example.com" --max-depth 2 --limit 50 --json
# Filter to specific paths
tvly crawl "https://example.com" --select-paths "/api/.*,/guides/.*" --exclude-paths "/blog/.*" --json
# Semantic focus (returns relevant chunks, not full pages)
tvly crawl "https://docs.example.com" --instructions "Find authentication docs" --chunks-per-source 3 --json| Option | Description |
|---|---|
--max-depth | Levels deep (1-5, default: 1) |
--max-breadth | Links per page (default: 20) |
--limit | Total pages cap (default: 50) |
--instructions | Natural language guidance for semantic focus |
--chunks-per-source | Chunks per page (1-5, requires --instructions) |
--extract-depth | basic (default) or advanced |
--format | markdown (default) or text |
--select-paths | Comma-separated regex patterns to include |
--exclude-paths | Comma-separated regex patterns to exclude |
--select-domains | Comma-separated regex for domains to include |
--exclude-domains | Comma-separated regex for domains to exclude |
--allow-external / --no-external | Include external links (default: allow) |
--include-images | Include images |
--timeout | Max wait (10-150 seconds) |
-o, --output | Save JSON output to file |
--output-dir | Save each page as a .md file in directory |
--json | Structured JSON output |
For agentic use (feeding results to an LLM):
Always use --instructions + --chunks-per-source. Returns only relevant chunks instead of full pages — prevents context explosion.
tvly crawl "https://docs.example.com" --instructions "API authentication" --chunks-per-source 3 --jsonFor data collection (saving to files):
Use --output-dir without --chunks-per-source to get full pages as markdown files.
tvly crawl "https://docs.example.com" --max-depth 2 --output-dir ./docs/--max-depth 1, --limit 20 — and scale up.--select-paths to focus on the section you need.--limit to prevent runaway crawls.Crawl websites and extract content from multiple pages via the Tavily CLI. Use this skill when the user wants to crawl a site, download documentation, extract an entire docs section, bulk-extract pages, save a site as local markdown files, or says "crawl", "get all the pages", "download the docs", "extract everything under /docs", "bulk extract", or needs content from many pages on the same domain. Supports depth/breadth control, path filtering, semantic instructions, and saving each page as a local markdown file.
The verbatim description from this skill’s front matter — the string an agent matches on to decide whether to load it.
Bash(tvly*)skills/tavily-crawl/SKILL.mdmain, last pushed 4 September 2026.SKILL.md, not by matching a directory convention. One layout observed: skills/*/SKILL.md.h1 and no skipped levels:.claude-plugin/marketplace.json by Tavily, declaring 1 plugin. It is read for editorial metadata only — never as the skill index, which is always the repository tree./tavily-ai/skills.md, and each skill at its own .md URL.