Skip to main content
The Scrape feature converts any URL into LLM-ready data formats including markdown, HTML, screenshots, and structured JSON. It handles JavaScript rendering, dynamic content, and can interact with pages before extracting data.

When to Use Scrape

Use Scrape when you need to:
  • Extract content from a single URL
  • Convert web pages to clean markdown for LLM processing
  • Capture screenshots of pages
  • Extract structured data using a schema
  • Interact with pages (login, click, scroll) before scraping
  • Extract brand identity (colors, fonts, typography)

Basic Usage

Response

Available Formats

You can request multiple formats in a single scrape:
  • markdown - Clean markdown content
  • html - Cleaned HTML
  • rawHtml - Original HTML
  • screenshot - Base64 encoded screenshot
  • links - All links found on the page
  • json - Structured data extraction
  • branding - Brand identity (colors, fonts, typography)

Get a Screenshot

Extract Brand Identity

Extract Structured Data

Extract structured data using a schema with JSON mode:
Output:

Extract with Prompt (No Schema)

You can also extract data using just a prompt without defining a schema:

Actions: Interact Before Scraping

Perform actions like clicking, typing, scrolling, and waiting before extracting content:
Actions are executed sequentially in the order specified. Use the wait action to allow dynamic content to load after interactions.

Best Practices

  • Request only the formats you need to minimize API usage
  • Use screenshot format to verify page rendering
  • For structured extraction, define a clear schema for consistent results
  • Use actions to handle pages requiring authentication or interaction
  • The branding format is useful for design research and competitor analysis

Next Steps

  • Learn about Batch Scraping to scrape multiple URLs efficiently
  • Use Crawl to scrape entire websites
  • Try Agent for autonomous data gathering