Anthropic has released the next generation of Claude models, Claude Opus 4 and Claude Sonnet 4, establishing new benchmarks for coding, advanced reasoning, and AI agent workflows. These hybrid models offer both near-instant responses and an "extended thinking" mode for deeper reasoning, alongside new capabilities such as parallel tool execution and improved memory.

  • Claude Opus 4 leads on SWE-bench (72.5%) and Terminal-bench (43.2%), sustaining performance on long-running tasks for several hours.
  • Claude Sonnet 4 achieves a state-of-the-art 72.7% on SWE-bench, offering an optimal mix of capability and efficiency as an upgrade to Sonnet 3.7.
  • Both models support extended thinking with tool use in beta, allowing them to alternate between reasoning and actions like web search.
  • Claude Code is now generally available with native integrations for VS Code and JetBrains, plus a new SDK for building custom agents.
  • New API capabilities include the code execution tool, MCP connector, Files API, and prompt caching for up to one hour.

The models are designed to advance AI strategies by pushing boundaries in coding and research while bringing frontier performance to everyday use cases through enhanced steerability and reduced shortcut behaviors.