← back to the directory
Library / SDK Retrieval & Memory

Unstructured

1Converts 2many 3file 4types 5to 6structure

the six Ws · specification

W1 Who

Data engineers building ETL pipelines that feed unstructured documents into vector databases and LLMs.

W2 What

Unstructured is an open-source ETL library, with an optional enterprise platform, that partitions PDFs, images, and 60+ file types into structured elements for chunking and embedding.

W3 Where

Self-hosted as a Python library or REST API; the enterprise Platform product is hosted SaaS.

W4 When

Widely adopted since 2022, with the enterprise Platform product actively expanded through 2026.

W5 Why

It exists to standardize messy, varied document formats into clean structured data for downstream AI workflows.

W6 With

Integrates with LangChain, LlamaIndex, and vector databases as a preprocessing step.

W7 Watch

Open source core (Apache-2.0), with a paid hosted Platform; self-hosted library needs no API key, hosted Platform requires an API key and sends documents to Unstructured's cloud.

ETL for LLMsdocument partitioning60+ file typeschunking and embedding

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.