ShieldGemma
the six Ws · specification
Google released ShieldGemma as a safety-classification model built on Gemma 2.
A set of models (2B, 9B, 27B) that classify text for harassment, hate speech, sexually explicit, and dangerous content.
Published on Hugging Face and Kaggle under the Google organization.
ShieldGemma was released in July 2024.
It gave developers an open, customizable content-moderation layer instead of relying only on closed moderation APIs.
Runs via transformers on top of the Gemma 2 base models, used as a classifier before or after generation.
Open-weight release by Google under the Gemma license, permitting broad research and commercial use with usage restrictions.