← back to the directory
Model Inference & Serving

Molmo

1Fully 2open 3multimodal 4models 5from 6Ai2

the six Ws · specification

W1 Who

The Allen Institute for AI (Ai2) built the Molmo family of multimodal models.

W2 What

Open vision-language models trained entirely on open datasets, including the PixMo image-caption corpus.

W3 Where

Published on Hugging Face under the allenai organization.

W4 When

Molmo was released in September 2024.

W5 Why

It aimed to make state-of-the-art multimodal AI fully open, including data, unlike most closed VLM competitors.

W6 With

Runs on PyTorch and transformers, built on Qwen2 and OLMo language model backbones with a CLIP-based vision encoder.

W7 Watch

Fully open release (weights, code, and training data) by Ai2 under Apache 2.0 license.

vision-languagemultimodalopen-sourceallen-ai

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.