← back to the directory
Library / SDK Retrieval & MemoryGuardrails & Structure

PyMuPDF4LLM

1Converts 2PDFs 3into 4markdown 5for 6RAG

the six Ws · specification

W1 Who

Developers building RAG pipelines who need clean text and structure extracted from PDFs and other documents.

W2 What

PyMuPDF4LLM converts PDFs and other documents into structured Markdown, JSON, or plain text optimized for RAG and LLM ingestion, including table and layout handling.

W3 Where

Used as a Python library built on PyMuPDF, running entirely locally without requiring GPU or cloud services.

W4 When

Released around 2024 by Artifex Software and actively maintained as of 2026.

W5 Why

Preserves reading order, tables, and layout from complex PDFs so downstream LLMs get cleaner context than raw text extraction.

W6 With

Built directly on the PyMuPDF library and integrates with frameworks like LlamaIndex and LangChain.

W7 Watch

Open source under GNU AGPLv3 with a paid commercial license available from Artifex Software; runs entirely locally with no data sent externally.

PDF parsingmarkdown conversionOCRlocal processingRAG data prep

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.