Cognition has released SWE-2, its most capable coding model to date, which is post-trained with reinforcement learning from Moonshot AI's 2.8T-parameter open model, Kimi K3. The model scores 50.0% on FrontierCode 1.1 Main, placing it within one point of Fable 5.1 while operating at 64% lower cost.

  • SWE-2 is Cognition's first model with selectable reasoning-effort levels, all trained in a single RL run using Pareto-informed cost penalties.
  • It achieves a score of 73.0% on DeepSWE 1.1 and 92.8% on Terminal-Bench 2.1, beating its Kimi K3 base on every row.
  • SWE-2 medium reduces turns by 58% and costs by 81% compared to SWE-1.7 on FrontierCode, while making its first real edit after a median of 18 steps versus 48 for the previous version.
  • The model is not available as open weights or via standalone API; it runs exclusively inside Devin Desktop, CLI, and upcoming Web/Fusion versions.

The release demonstrates that scaling RL to the multi-trillion-parameter regime allows Cognition to find substantial headroom on top of K3, adding 5 to 6 points on many benchmarks.