GAIR-NLP introduces LIMO, a model demonstrating that complex mathematical reasoning can be effectively elicited using only 817 curated training samples. This approach challenges the conventional assumption that supervised fine-tuning requires massive datasets to avoid memorization.

  • LIMO achieves 57.1% accuracy on the AIME benchmark and 94.8% on MATH.
  • The model uses just 1% of the training data required by previous strong SFT-based models.
  • It demonstrates exceptional out-of-distribution generalization, improving performance by 40.5% absolute across 10 diverse benchmarks.
  • Results outperform models trained on 100x more data, supporting the "Less-Is-More Reasoning Hypothesis."

The authors propose that sophisticated reasoning capabilities emerge when foundation models leverage rich pre-trained knowledge through minimal, precisely orchestrated cognitive templates.