Skip to main content
Batch Scrape allows you to scrape multiple URLs efficiently with parallel processing. It’s ideal when you have a list of specific URLs to scrape and want to process them all at once.

When to Use Batch Scrape

Use Batch Scrape when you need to:
  • Scrape a known list of URLs
  • Process hundreds or thousands of pages in parallel
  • Extract data from multiple pages with the same structure
  • Update data from a list of product pages, articles, or profiles
  • Scrape URLs discovered from Map or other sources

Basic Usage

The SDKs automatically wait for batch scraping to complete. For manual control, use the async methods below.

Asynchronous Batch Scrape

For large batches, start the job asynchronously and poll for status:

Batch Scrape with Structured Data

Extract structured data from multiple pages:

Real-Time Updates

Receive updates as pages are scraped:

Manual Pagination

For very large batches, you can manually paginate through results:

Webhooks

Receive notifications when pages are scraped:

Webhook Events

  • batch_scrape.started - Job has started
  • batch_scrape.page - A page has been scraped
  • batch_scrape.completed - All pages scraped successfully
  • batch_scrape.failed - Job failed

Handling Invalid URLs

Ignore invalid URLs instead of failing the entire batch:
Invalid URLs will be returned in the invalidURLs field of the response.

Cancel a Batch Job

Use Cases

Scrape URLs from Map

1

Map the website

2

Batch scrape discovered URLs

Update Product Catalog

Monitor Competitor Prices

Best Practices

  • Use Batch Scrape for lists of known URLs, not for discovering URLs
  • Combine with Map to discover URLs first, then batch scrape them
  • Use webhooks for very large batches to avoid polling
  • Set ignoreInvalidURLs: true when working with uncertain URL lists
  • Request only the formats you need to minimize processing time
  • Use structured extraction with schemas for consistent data
  • Consider rate limits and costs when batch scraping thousands of URLs

Limits and Performance

  • URLs are processed in parallel for maximum speed
  • No hard limit on the number of URLs per batch
  • Each URL counts as one scrape credit
  • Failed URLs can be retried individually

Next Steps

  • Learn about Map to discover URLs for batch scraping
  • Use Crawl when you need to scrape an entire site
  • Try Scrape for individual URLs with more control