← back to the directory
Library / SDK Automation & IntegrationRetrieval & Memory

Crawlee

1Reliable 2web 3scraping 4library 5for 6LLMs

the six Ws · specification

W1 Who

Developers building crawlers or data pipelines that feed scraped web content into AI and RAG systems.

W2 What

A Node.js and TypeScript library for web scraping and browser automation that unifies Playwright, Puppeteer, Cheerio, and raw HTTP crawling with built-in proxy rotation and queue management.

W3 Where

Runs as a library inside Node.js scraping jobs, either headless or headful, locally or in the cloud.

W4 When

Use when a project needs dependable, resumable crawls to collect training data or documents for retrieval.

W5 Why

It handles the retry, queueing, and anti-blocking plumbing that hand-rolled scrapers usually get wrong.

W6 With

Wraps Playwright, Puppeteer, Cheerio, and JSDOM behind one consistent crawler API.

W7 Watch

Open source, Apache 2.0 licensed, maintained by Apify; free to use with optional paid Apify cloud hosting.

web scrapingbrowser automationtypescriptproxy rotationdata extraction for llms

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.