Sangfor has deployed SESG (Self-Evolving Safety Guardrails), a multi-agent system that continuously adapts LLM safety defenses to new threats in production. The system monitors live traffic for novel jailbreaks and harmful categories, then automatically synthesizes training data and updates the model without manual intervention.

  • SESG uses generation, validation, and routing agents to diagnose failures and create targeted training batches.
  • A 1.7B guardrail adapts to new threats in 16-24 hours with only about 2 hours of human effort, compared to 40-90 hours for manual processes.
  • The system outperforms static guardrails ranging from 0.6B to 9B parameters and an adaptive baseline on six emerging threats.
  • Since April 2026, SESG has autonomously closed 14 of 15 new threat scenarios in two months.

This approach allows safety models to keep pace with rapidly evolving attack vectors while maintaining general screening competence.