Skip to main content
Web Scrape stage showing content extraction from URLs
The Web Scrape stage fetches content from web URLs using the Engine’s Playwright service. It handles JavaScript rendering and main-content extraction to add web content to your retrieval pipeline.
Stage Category: APPLY (Enriches documents with scraped content)Transformation: N documents → N documents (with web content added)

When to Use

When NOT to Use

Parameters

Configuration Examples

Output Schema

Default Output

The extracted content is stored as an object at output_field.

With Raw HTML

Set include_html to true to add the page HTML to the same object.

Failed Fetch

With the default on_error: skip, a document whose fetch fails passes through unchanged, without the output field.

Scraping Strategies

Use strategy: "javascript" when scraping JavaScript-heavy sites, and raise timeout to 20-30 seconds.

Performance

Web scraping adds significant latency. Use sparingly and consider pre-indexing frequently accessed content.

Common Pipeline Patterns

Enrich Documents with Referenced Content

Scrape and Summarize

Error Handling

on_error accepts three values. skip (default) passes the document through unchanged. remove drops the document from the results. raise fails the stage.

Rate Limits and Best Practices

  1. Batch wisely: Limit to 5-10 URLs per pipeline run
  2. Cache results: Consider storing scraped content
  3. Use timeouts: Set an appropriate timeout (in seconds) for your use case