← back to the directory
Library / SDK Retrieval & MemoryGuardrails & Structure

Sycamore

1AI 2powered 3ETL 4engine 5for 6documents

the six Ws · specification

W1 Who

Data engineers building RAG and analytics pipelines over complex unstructured documents like PDFs and presentations.

W2 What

Sycamore is an open source document processing engine from Aryn that segments, OCRs, and extracts tables and images from documents for ETL, RAG, and analytics.

W3 Where

Used as a Python library, optionally paired with the Aryn DocParse GPU powered API, loading results into vector databases like OpenSearch or Pinecone.

W4 When

Released by Aryn around 2023 and actively maintained as of 2026.

W5 Why

Handles messy real world documents with embedded tables and figures more reliably than simple text extraction, improving downstream RAG quality.

W6 With

Can run fully locally or depend on the hosted Aryn DocParse API for GPU accelerated segmentation and OCR.

W7 Watch

Open source under Apache-2.0 for the core library; the optional Aryn DocParse component is a hosted API that requires credentials and sends document data to Aryn's servers.

document ETLPDF parsingRAG data preptable extractionOCR

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.