Skills
Skill 5 of 8
Extract clean markdown or text content from specific URLs via the Tavily CLI.
1 minute · 291 words · 7 sections
Install
npx skills add tavily-ai/skills --skill tavily-extractnpx skills add tavily-ai/skills/plugin marketplace add tavily-ai/skillsThe first command installs just this skill, by the name in its SKILL.md; the second installs the whole repository.
Extract clean markdown or text content from one or more URLs.
Run extract directly when tvly is available. Extract supports capped keyless
access, so do not look for an API key or authenticate before the first request.
If tvly is missing, follow the tavily-cli setup
before retrying. If the keyless cap is reached in an interactive session, run
tvly login to open browser OAuth, then retry the original extraction once. In
an unattended environment, report the cap and authentication options instead
of starting an interactive flow. Do not start a second login immediately after
guided setup has completed.
# Single URL
tvly extract "https://example.com/article" --json
# Multiple URLs
tvly extract "https://example.com/page1" "https://example.com/page2" --json
# Query-focused extraction (returns relevant chunks only)
tvly extract "https://example.com/docs" --query "authentication API" --chunks-per-source 3 --json
# JS-heavy pages
tvly extract "https://app.example.com" --extract-depth advanced --json
# Save to file
tvly extract "https://example.com/article" -o article.json| Option | Description |
|---|---|
--query | Rerank chunks by relevance to this query |
--chunks-per-source | Chunks per URL (1-5, requires --query) |
--extract-depth | basic (default) or advanced (for JS pages) |
--format | markdown (default) or text |
--include-images | Include image URLs |
--timeout | Max wait time (1-60 seconds) |
-o, --output | Save the JSON response to a file |
--json | Structured JSON output |
| Depth | When to use |
|---|---|
basic | Simple pages, fast — try this first |
advanced | JS-rendered SPAs, dynamic content, tables |
--query + --chunks-per-source to get only relevant content instead of full pages.basic first, fall back to advanced if content is missing.--timeout for slow pages (up to 60s).failed_results even after exit code 0. A successful request can
still return no extracted pages. Retry the affected URL with advanced when
appropriate, otherwise report the per-URL failure instead of treating the
request as complete.--include-raw-content), skip the extract step.Extract clean markdown or text content from specific URLs via the Tavily CLI. Use this skill when the user has one or more URLs and wants their content, says "extract", "grab the content from", "pull the text from", "get the page at", "read this webpage", or needs clean text from web pages. Handles JavaScript-rendered pages, returns LLM-optimized markdown, and supports query-focused chunking for targeted extraction. Can process up to 20 URLs in a single call.
The verbatim description from this skill’s front matter — the string an agent matches on to decide whether to load it.
Bash(tvly*)skills/tavily-extract/SKILL.mdmain, last pushed 4 September 2026.SKILL.md, not by matching a directory convention. One layout observed: skills/*/SKILL.md.h1 and no skipped levels:.claude-plugin/marketplace.json by Tavily, declaring 1 plugin. It is read for editorial metadata only — never as the skill index, which is always the repository tree./tavily-ai/skills.md, and each skill at its own .md URL.