started · updated
Mistral AI releases Shieldstral, 3‑B parameter AI safety model
Mistral AI has launched Shieldstran, a 3‑billion‑parameter open‑weight guard model that classifies text, images and mixed content for safety. The model runs on a single 16 GB GPU and is released under the Apache 2.0 licence, allowing on‑premise deployment.
In Mistral’s internal benchmarks Shieldstran achieved an overall F1 score of 84.9 % on 13 text‑safety tests, matching the performance of 20‑billion‑parameter models such as GPT‑OSS Safeguard. It also scored 83.8 % on multimodal safety, 91.3 % on policy‑adaptability, and 91.5 % on refusal detection. The model answers safety queries in a question‑and‑answer format, letting developers supply plain‑language policies at inference time without additional training.
Shieldstran’s small size and low hardware requirements aim to reduce the cost of adding a moderation layer to generative‑AI applications while maintaining high safety standards.