Ramalama
the six Ws · specification
Developers who prefer container native workflows for running and serving local models.
RamaLama is an open source tool from the containers.org ecosystem that pulls models as OCI artifacts and runs them for inference inside rootless Podman or Docker containers.
Runs locally via a CLI on Linux, macOS, or Windows, pulling models from Hugging Face, Ollama, or OCI registries.
Launched in 2024 by the Podman and containers community and updated regularly.
Applies standard container tooling and isolation to model serving so AI workloads fit existing DevOps pipelines.
Built on Podman or Docker and uses llama.cpp or vLLM as the underlying inference engine.
Open source under MIT on GitHub, no account required, runs fully offline once images are pulled.