Available Formats
Specify the formats you want using theformats parameter:
from firecrawl import Firecrawl
app = Firecrawl(api_key="fc-YOUR_API_KEY")
doc = app.scrape(
url="https://firecrawl.dev",
formats=["markdown", "html", "screenshot"]
)
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
const doc = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['markdown', 'html', 'screenshot']
});
curl -X POST 'https://api.firecrawl.dev/v2/scrape' \
-H 'Authorization: Bearer fc-YOUR_API_KEY' \
-H 'Content-Type: application/json' \
-d '{
"url": "https://firecrawl.dev",
"formats": ["markdown", "html", "screenshot"]
}'
Markdown
Clean, LLM-ready markdown format. This is the default format and is ideal for feeding content into AI models.from firecrawl import Firecrawl
app = Firecrawl(api_key="fc-YOUR_API_KEY")
doc = app.scrape("https://firecrawl.dev", formats=["markdown"])
print(doc.markdown)
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
const doc = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['markdown']
});
console.log(doc.markdown);
{
"success": true,
"data": {
"markdown": "# Firecrawl Docs\n\nTurn websites into LLM-ready data...",
"metadata": {
"title": "Quickstart | Firecrawl",
"description": "Firecrawl allows you to turn entire websites into LLM-ready markdown",
"sourceURL": "https://docs.firecrawl.dev",
"statusCode": 200
}
}
}
Markdown format automatically removes headers, footers, navigation, and other non-main content when
onlyMainContent is true (default).HTML
Cleaned HTML version of the page content.from firecrawl import Firecrawl
app = Firecrawl(api_key="fc-YOUR_API_KEY")
doc = app.scrape("https://firecrawl.dev", formats=["html"])
print(doc.html)
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
const doc = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['html']
});
console.log(doc.html);
Raw HTML
The complete, unmodified HTML of the page including all scripts, styles, and metadata.from firecrawl import Firecrawl
app = Firecrawl(api_key="fc-YOUR_API_KEY")
doc = app.scrape("https://firecrawl.dev", formats=["rawHtml"])
print(doc.raw_html)
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
const doc = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['rawHtml']
});
console.log(doc.rawHtml);
Raw HTML includes everything on the page and can be very large. Use this only when you need the complete page source.
Links
Extract all links found on the page.from firecrawl import Firecrawl
app = Firecrawl(api_key="fc-YOUR_API_KEY")
doc = app.scrape("https://firecrawl.dev", formats=["links"])
for link in doc.links:
print(link)
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
const doc = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['links']
});
doc.links.forEach(link => console.log(link));
{
"success": true,
"data": {
"links": [
"https://firecrawl.dev/pricing",
"https://firecrawl.dev/blog",
"https://docs.firecrawl.dev"
]
}
}
Screenshot
Capture a screenshot of the page. Screenshots are returned as base64-encoded images.from firecrawl import Firecrawl
import base64
app = Firecrawl(api_key="fc-YOUR_API_KEY")
# Viewport screenshot
doc = app.scrape("https://firecrawl.dev", formats=["screenshot"])
print(doc.screenshot) # Base64 encoded image
# Full page screenshot
doc = app.scrape("https://firecrawl.dev", formats=["screenshot@fullPage"])
print(doc.screenshot)
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
// Viewport screenshot
const doc = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['screenshot']
});
console.log(doc.screenshot); // Base64 encoded image
// Full page screenshot
const doc2 = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['screenshot@fullPage']
});
console.log(doc2.screenshot);
Use
screenshot for viewport-sized screenshots or screenshot@fullPage to capture the entire page including content below the fold.JSON (Structured Extraction)
Extract structured data from pages using AI with a schema or prompt.With Schema
Define a precise structure for the data you want to extract:from firecrawl import Firecrawl
from pydantic import BaseModel
app = Firecrawl(api_key="fc-YOUR_API_KEY")
class CompanyInfo(BaseModel):
company_mission: str
is_open_source: bool
is_in_yc: bool
result = app.scrape(
'https://firecrawl.dev',
formats=[{"type": "json", "schema": CompanyInfo.model_json_schema()}]
)
print(result.json)
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
const schema = {
type: 'object',
properties: {
company_mission: { type: 'string' },
is_open_source: { type: 'boolean' },
is_in_yc: { type: 'boolean' }
},
required: ['company_mission', 'is_open_source', 'is_in_yc']
};
const result = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['markdown'],
jsonOptions: { schema }
});
console.log(result.json);
{
"success": true,
"data": {
"json": {
"company_mission": "Turn websites into LLM-ready data",
"is_open_source": true,
"is_in_yc": true
}
}
}
With Prompt (No Schema)
Extract data using natural language without defining a strict schema:from firecrawl import Firecrawl
app = Firecrawl(api_key="fc-YOUR_API_KEY")
result = app.scrape(
'https://firecrawl.dev',
formats=[{"type": "json", "prompt": "Extract the company mission"}]
)
print(result.json)
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
const result = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['markdown'],
jsonOptions: {
prompt: 'Extract the company mission'
}
});
console.log(result.json);
Branding
Extract brand identity information including colors, fonts, and typography.from firecrawl import Firecrawl
app = Firecrawl(api_key="fc-YOUR_API_KEY")
doc = app.scrape("https://firecrawl.dev", formats=["branding"])
print(doc.branding)
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
const doc = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['branding']
});
console.log(doc.branding);
{
"success": true,
"data": {
"branding": {
"logo": "https://firecrawl.dev/logo.png",
"colors": {
"primary": "#FF6B35",
"secondary": "#004E89"
},
"fonts": [
{"family": "Inter"},
{"family": "Roboto"}
],
"typography": {
"headingFont": "Inter",
"bodyFont": "Roboto"
}
}
}
}
Branding extraction derives information by executing on-page JavaScript to analyze computed styles and detect brand assets.
Change Tracking
Track changes to web pages over time using thechangeTracking format.
from firecrawl import Firecrawl
app = Firecrawl(api_key="fc-YOUR_API_KEY")
doc = app.scrape(
url="https://example.com",
formats=["markdown", "changeTracking"],
change_tracking_options={
"modes": ["git-diff"]
}
)
if doc.change_tracking:
print(f"Status: {doc.change_tracking['changeStatus']}")
if doc.change_tracking.get('diff'):
print(doc.change_tracking['diff'])
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
const doc = await app.scrape({
url: 'https://example.com',
formats: ['markdown', 'changeTracking'],
changeTrackingOptions: {
modes: ['git-diff']
}
});
if (doc.changeTracking) {
console.log(`Status: ${doc.changeTracking.changeStatus}`);
if (doc.changeTracking.diff) {
console.log(doc.changeTracking.diff);
}
}
Change tracking requires the
markdown format to also be specified. The first scrape establishes a baseline for future comparisons.Combining Multiple Formats
You can request multiple formats in a single scrape to get different views of the same content:from firecrawl import Firecrawl
app = Firecrawl(api_key="fc-YOUR_API_KEY")
doc = app.scrape(
url="https://firecrawl.dev",
formats=["markdown", "html", "screenshot", "links", "branding"]
)
print(f"Title: {doc.metadata['title']}")
print(f"Links found: {len(doc.links)}")
print(f"Primary color: {doc.branding['colors']['primary']}")
print(f"Content length: {len(doc.markdown)} chars")
import Firecrawl from '@mendable/firecrawl-js';
const app = new Firecrawl({ apiKey: 'fc-YOUR_API_KEY' });
const doc = await app.scrape({
url: 'https://firecrawl.dev',
formats: ['markdown', 'html', 'screenshot', 'links', 'branding']
});
console.log(`Title: ${doc.metadata.title}`);
console.log(`Links found: ${doc.links.length}`);
console.log(`Primary color: ${doc.branding.colors.primary}`);
console.log(`Content length: ${doc.markdown.length} chars`);
Best Practices
Request only what you need: Each format adds to processing time and response size. Only request formats you’ll actually use.
Use JSON extraction for structured data: Instead of parsing markdown or HTML yourself, use JSON format with a schema to extract exactly what you need.
Combine formats strategically: Request
markdown + links for content analysis and navigation discovery, or screenshot + branding for visual analysis.