tech
Introducing Shieldstral.
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining. Released under Apache 2.0, it delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GPU.
TL;DR
- Shieldstral is a 3B open-weights multimodal safety classifier.
- It frames content moderation as a policy-adaptive question-answering task, accepting plain-language policies at inference time.
- The classifier unifies text and image safety evaluation without retraining.
- It outperforms models up to 7x its size on text safety and sets a new state-of-the-art on multimodal moderation.
- Shieldstral delivers calibrated safety scores across diverse benchmarks and runs efficiently on a single 16GB NVIDIA GPU.
- It is released under Apache 2.0 license.