Skip to main content
The Crawl feature allows you to scrape all pages of a website with a single request. It automatically discovers URLs, respects your limits and filters, and scrapes each page according to your specifications.

When to Use Crawl

Use Crawl when you need to:
  • Extract content from an entire website or documentation site
  • Build a knowledge base from web content
  • Index website content for search
  • Monitor website changes over time
  • Create datasets from multi-page websites

Basic Usage

1

Start a Crawl

2

Response

The initial response contains a job ID:
3

Check Status (if using async)

Status Response:
The SDKs handle polling automatically, waiting for the crawl to complete before returning results.

Asynchronous Crawling

For long-running crawls, start the job asynchronously and poll for status:

Crawl Options

Control Crawl Depth

Include/Exclude Paths

Use regex patterns to filter URLs:
  • allowBackwardLinks: false - Only crawls deeper (child) URLs
  • allowBackwardLinks: true - Crawls any internal links, including siblings and parents
  • allowExternalLinks: true - Follows links to external websites

Ignore Sitemap

Real-Time Updates with WebSockets

Get updates as pages are crawled:

Cancel a Crawl

Best Practices

  • Use limit to control costs and crawl time
  • Set appropriate maxDepth to avoid crawling too deep
  • Use includePaths and excludePaths to focus on relevant content
  • Enable WebSocket watchers for real-time monitoring of large crawls
  • Set a delay between scrapes to respect website rate limits
  • Start with a small limit to test your configuration before scaling up

Next Steps

  • Learn about Map to discover URLs before crawling
  • Use Batch Scrape for known lists of URLs
  • Try Search to find and scrape search results