Stage Category: APPLY (Enriches documents with scraped content)Transformation: N documents → N documents (with web content added)
When to Use
When NOT to Use
Parameters
Configuration Examples
Output Schema
Default Output
The extracted content is stored as an object atoutput_field.
With Raw HTML
Setinclude_html to true to add the page HTML to the same object.
Failed Fetch
With the defaulton_error: skip, a document whose fetch fails passes through unchanged, without the output field.
Scraping Strategies
Performance
Common Pipeline Patterns
Enrich Documents with Referenced Content
Scrape and Summarize
Error Handling
on_error accepts three values. skip (default) passes the document through unchanged. remove drops the document from the results. raise fails the stage.
Rate Limits and Best Practices
- Batch wisely: Limit to 5-10 URLs per pipeline run
- Cache results: Consider storing scraped content
- Use timeouts: Set an appropriate
timeout(in seconds) for your use case
Related
- External Web Search - Search the web (Exa)
- API Call - General HTTP enrichment
- Document Enrich - Collection joins

