Claude Sonnet 5 Is Live: Real Agentic Gains at a Price Built for Daily Use
The exact benchmark numbers separating Sonnet 5 from Sonnet 4.6 and Opus 4.8, and what they mean for real work
What matters today
The exact benchmark numbers separating Sonnet 5 from Sonnet 4.6 and Opus 4.8, and what they mean for real work
Article roadmap
What you will learn
-
The exact benchmark numbers separating Sonnet 5 from Sonnet 4.6 and Opus 4.8, and what they mean for real work
-
Where Sonnet 5 runs today: claude.ai, Claude Code, AWS Bedrock, Google Vertex AI, and Microsoft Foundry
-
A clear rule for choosing Sonnet 5 over Opus 4.8, with the one setting that controls both cost and quality
-
A step-by-step contract review workflow you can run this week, plus a second use case for competitive research
-
The pricing catch buried in Anthropic's footnotes that changes how much you actually pay per task
Every few months, a new model launches with a press release full of "significant improvements" and no numbers you can act on. Sonnet 5, which Anthropic shipped on June 30, 2026, is not that. It comes with a benchmark table, an effort dial that trades cost for quality in real time, and a price that stays flat instead of climbing.
The decision in front of Executives is not whether to use Claude. It is which model to route each job to, and whether paying Opus prices is still necessary for tasks that used to require it. Get that wrong and you either overpay for routine work or undershoot on tasks that need real reasoning depth.
Below: the numbers behind the "agentic" claim, where the model runs, when to reach for Opus 4.8 instead, and a full walkthrough of a contract review workflow built for Sonnet 5's new effort levels.
What Actually Changed
"Agentic" gets thrown around a lot. In practice it means a model that plans a multi-step task, uses tools like a browser or terminal, checks its own work, and keeps going without you re-prompting it after every step. Anthropic's launch post describes testers watching Sonnet 5 finish jobs "where previous Sonnet models would stop short," including a case where the model investigated a bug, wrote a test to reproduce it, fixed it, then stashed the fix to confirm the bug returned without it, all in a single unprompted pass.
The clearest apples-to-apples read sits in SWE-bench Pro, a real-world software engineering benchmark. It is the only benchmark in Anthropic's release notes with a published score for Sonnet 4.6, Sonnet 5, and Opus 4.8 alike, so it is the one comparison below that is not missing a data point for any of the three models.
Free subscriber access
Keep reading with your free PromptHacker subscription.
Enter the email on your PromptHacker subscription. If access is available, a secure link will arrive in that inbox.
New to PromptHacker? Join the free weekly briefing first.
Three deep dives. Four useful moves. One email worth opening.
PromptHacker turns the AI firehose into practical next steps for work, health, family, and everything time keeps trying to steal.