The report introduces Zing, an integrated framework designed to enhance the social intelligence of large language models by measuring, internalizing, and grounding social capabilities. The authors present SoMBench, a psychology-grounded benchmark spanning 17 secondary dimensions and 3,481 expert-verified instances, which reveals that current LLMs have significant room for improvement with no model reaching near-ceiling performance.

  • Zing employs a diagnosis-driven training recipe combining supervised fine-tuning, on-policy distillation, and rubric-based reinforcement learning to internalize social skills.
  • Zing-27B-Stage2 achieves the best average score across five social-cognition benchmarks, while Zing-32B-Stage2 remains competitive with DeepSeek-V4-Pro.
  • Actio is introduced as a harness-controlled inference architecture that routes four typed supports—PRISM, Starling, SAGE, and gated RAG—into reasoning.
  • The full Actio harness improves 14 of 15 model-benchmark pairs across five base models and three benchmarks, demonstrating the effectiveness of typed runtime support.

These results demonstrate that socially intelligent LLMs require coordinated advances in evaluation, parametric internalization, and deployment-time grounding to effectively infer mental states and adapt behavior.