Mistral Releases Shieldstral, a Small Open-Weights Safety Classifier
Mistral released Shieldstral, a 3B multimodal classifier that applies plain-language policies at inference time.
Mistral released Shieldstral, a 3B open-weights multimodal safety classifier under Apache 2.0. The company said the model could evaluate text and images against plain-language policies at inference time without retraining.
The release focused on a practical layer of the agent stack: policy enforcement that can adapt to a product's audience and rules. Smaller classifiers can be useful when teams need configurable moderation closer to the application, rather than relying only on a large model's general behavior.
Editorial sources
Every claim in this briefing traces back to the references below.
- Introducing Shieldstral — Mistral AI — Primary model announcement https://mistral.ai/news/shieldstral/