May 2025
Claude Sonnet 4 and Opus 4 — Anthropic's Most Capable Generation
Anthropic released Claude Sonnet 4 and Claude Opus 4, with Sonnet 4 achieving 72.7% on SWE-bench Verified — a new record at launch — and Opus 4 positioned as the highest-capability option for the most demanding tasks. Both models featured improved instruction following, significantly reduced unnecessary refusals compared to earlier versions, and enhanced agentic capabilities for multi-step workflows. Claude Sonnet 4 became the primary model powering Cursor's AI coding assistant.
Why it mattersClaude Sonnet 4 and Opus 4 set a new bar for agentic task completion, meaning teams building multi-step automation should re-evaluate their model choice for reliability on long-horizon tasks.
Try itRun your existing Claude 3.7 agent benchmarks against Claude Sonnet 4 today — measure task completion rate and error recovery on the same 20 representative workflows.