Researchers introduce Seed1.5-Thinking, a Mixture-of-Experts model with 20B activated and 200B total parameters that reasons through thinking before responding. The method achieves 86.7 on AIME 2024, 55.0 on Codeforces, and 77.3 on GPQA, demonstrating strong reasoning abilities in STEM and coding.

  • Surpasses DeepSeek R1 by 8% in win rate on non-reasoning tasks.
  • Develops two internal benchmarks, BeyondAIME and Codeforces, for assessing generalized reasoning.

The model demonstrates notable generalization across diverse domains, indicating broader applicability beyond pure reasoning tasks.