tech

Introducing Shieldstral.

Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining. Released under Apache 2.0, it delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GPU.

Introducing Shieldstral.

TL;DR

  • Shieldstral is a 3B open-weights multimodal safety classifier.
  • It frames content moderation as a policy-adaptive question-answering task, accepting plain-language policies at inference time.
  • The classifier unifies text and image safety evaluation without retraining.
  • It outperforms models up to 7x its size on text safety and sets a new state-of-the-art on multimodal moderation.
  • Shieldstral delivers calibrated safety scores across diverse benchmarks and runs efficiently on a single 16GB NVIDIA GPU.
  • It is released under Apache 2.0 license.