← back to the directory
Library / SDK Fine-tuning & TrainingInference & Serving

SentencePiece

1Unsupervised 2subword 3tokenizer 4for 5language 6models

the six Ws · specification

W1 Who

Developed and maintained by Google's language research team.

W2 What

SentencePiece is an unsupervised, language independent subword tokenizer and detokenizer trained directly from raw text.

W3 Where

Hosted on GitHub at google/sentencepiece and distributed via PyPI, pip and C++ bindings.

W4 When

Released in 2018 and now in a stable, low churn maintenance phase.

W5 Why

It lets teams train a consistent subword vocabulary directly from raw text without language specific pre-tokenization rules, unlike earlier tokenizers.

W6 With

Implements BPE and unigram language model algorithms and is used by T5, ALBERT, Llama and many other models.

W7 Watch

Apache 2.0 licensed, open source, maintained by Google, roughly 11k GitHub stars, stable and mostly in maintenance mode with low commit velocity.

tokenizationsubwordunsupervisedlanguage-agnosticgoogle

for agents & scripts

Reading this as a machine? Query it directly.

Search is open JSON - no key. Report telemetry after using a tool and it feeds that tool’s Proof Score. Or speak MCP to /mcp and discover tools mid-loop.