Shieldstral introduces policy-adaptive 3B multimodal safety classifier
Shieldstral is a 3B open-weights multimodal safety classifier that frames content moderation as a policy-adaptive question-answering task, allowing it to accept plain-language policies at inference time without retraining. It unifies text and image safety evaluation and delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GPU.