When to Use Crawl
Use Crawl when you need to:- Extract content from an entire website or documentation site
- Build a knowledge base from web content
- Index website content for search
- Monitor website changes over time
- Create datasets from multi-page websites
Basic Usage
1
Start a Crawl
- Python
- JavaScript
- cURL
2
Response
The initial response contains a job ID:
3
Check Status (if using async)
- Python
- JavaScript
- cURL
The SDKs handle polling automatically, waiting for the crawl to complete before returning results.
Asynchronous Crawling
For long-running crawls, start the job asynchronously and poll for status:- Python
- JavaScript
Crawl Options
Control Crawl Depth
- Python
- JavaScript
Include/Exclude Paths
Use regex patterns to filter URLs:- Python
- JavaScript
Allow Backward and External Links
- Python
- JavaScript
allowBackwardLinks: false- Only crawls deeper (child) URLsallowBackwardLinks: true- Crawls any internal links, including siblings and parentsallowExternalLinks: true- Follows links to external websites
Ignore Sitemap
- Python
- JavaScript
Real-Time Updates with WebSockets
Get updates as pages are crawled:- Python
- JavaScript
Cancel a Crawl
- Python
- JavaScript
Best Practices
Next Steps
- Learn about Map to discover URLs before crawling
- Use Batch Scrape for known lists of URLs
- Try Search to find and scrape search results