← back to the directory
Library / SDK Retrieval & Memory

olmOCR

1Converts 2scanned 3PDFs 4into 5clean 6text

the six Ws · specification

W1 Who

The Allen Institute for AI, Ai2, released olmOCR.

W2 What

olmOCR is a document parsing toolkit built around a fine-tuned vision-language model that converts scanned PDFs and images into clean, structured plain text or markdown.

W3 Where

Hosted on GitHub at allenai/olmocr and on Hugging Face under allenai.

W4 When

Released in February 2025.

W5 Why

It was built to accurately and cheaply convert large PDF corpora into training-ready text for language model pretraining.

W6 With

Runs on the olmOCR-7B vision-language model served via vLLM or SGLang.

W7 Watch

Released under the Apache 2.0 license by the Allen Institute for AI; fully open weights and code with no gating.

OCRdocument parsingPDFAllen AI

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.