When to Use Batch Scrape
Use Batch Scrape when you need to:- Scrape a known list of URLs
- Process hundreds or thousands of pages in parallel
- Extract data from multiple pages with the same structure
- Update data from a list of product pages, articles, or profiles
- Scrape URLs discovered from Map or other sources
Basic Usage
- Python
- JavaScript
- cURL
The SDKs automatically wait for batch scraping to complete. For manual control, use the async methods below.
Asynchronous Batch Scrape
For large batches, start the job asynchronously and poll for status:- Python
- JavaScript
Batch Scrape with Structured Data
Extract structured data from multiple pages:- Python
- JavaScript
Real-Time Updates
Receive updates as pages are scraped:- JavaScript
Manual Pagination
For very large batches, you can manually paginate through results:- Python
Webhooks
Receive notifications when pages are scraped:- cURL
Webhook Events
batch_scrape.started- Job has startedbatch_scrape.page- A page has been scrapedbatch_scrape.completed- All pages scraped successfullybatch_scrape.failed- Job failed
Handling Invalid URLs
Ignore invalid URLs instead of failing the entire batch:- cURL
invalidURLs field of the response.
Cancel a Batch Job
- Python
- JavaScript
Use Cases
Scrape URLs from Map
1
Map the website
2
Batch scrape discovered URLs
Update Product Catalog
Monitor Competitor Prices
Best Practices
Limits and Performance
- URLs are processed in parallel for maximum speed
- No hard limit on the number of URLs per batch
- Each URL counts as one scrape credit
- Failed URLs can be retried individually