BentoML
the six Ws · specification
Maintained by BentoML (bentoml).
A framework for packaging models and AI apps into standardized services and serving them as APIs.
Runs self-hosted as a Python framework, deployable to containers, Kubernetes, or cloud.
Reach for it when productionizing and serving models, including LLMs, with scaling.
Standardizes build, packaging, and deployment with adaptive batching and flexibility.
Python and model artifacts; GPUs for compute-heavy models.
Built by BentoML; open-source (Apache-2.0) with a paid managed cloud (open-core); self-hostable.