Sonnet 5
Where does this model even fit? Maybe we can ask Mythos.
The much anticipated Claude Sonnet 5 was released today, and I'm really not even sure why. It's fairly early to make a full assessment of it, but there are some questions to be asked right away. Especially regarding the announcement itself.
First, I'm pretty sure they just released something to get everyone off their backs about Fable and Mythos. That's really what everyone wants. In their defense, a new Sonnet was kind of overdue, so why not.
Okay, let's get right into it. This isn't really an improvement on anything it seems. Sure, it's better than Sonnet 4.6, but that only matters in a world where there's no Opus. They gave us their benchmarks. Here's where it falls:
- almost as good as Opus 4.8
- better than Sonnet 4.6 (duh)
- same price-per-token as Sonnet 4.6 (cool)
On the surface, that all seems well and good... until you actually use it. In their (Anthropic) announcement, they shared some charts. And here's the first question: why would you include these charts? They make this look bad.
And, okay, here's some stream of consciousness in the middle of this post. THEY CHANGED THE CHART.
I went into their announcement to grab screenshots of the chart to make my point, and it's different than it was earlier, which could change this whole post.
Here's the current chart:

Here's the original (which I only have because I was ranting about this earlier):

The original chart basically shows that Sonnet 5 is worse than Opus 4.8 in cost-per-task across the board.
This is where the rubber meets the road. It doesn't matter if Sonnet is cheaper on a cost-per-token basis. To get the same output, it costs more money. ~bUt iT's a bEtTeR SoNnEt~ so what? if this was their frontier model, then fine, it's a pretty massive improvement. but it's not.
anyway, there's a new chart. this one pretty much shows that Sonnet 5 is identical to Opus 4.8, except that it's more expensive. I guess Sonnet 5 Medium is nice? Gives you Sonnet 4.6 capability at a fraction of the price? Maybe it's a convenience play. You can just use this single model at different effort levels and get a wide range of value out of it?
I'm curious how this benchmarks against Haiku.
There's some more in the announcement about it not being able to exploit, which is probably to appease the government? I suppose that's nice, but again, just makes the model look less capable.
Anyway, my question still stands: where does this fit?
I'm relatively new to actually using these models, but there are just a few different consumers.
- people that just benchmark models and see what they can do, track progress, create content around them, etc
- people that need the best frontier model, whatever it costs, to do complex work, architecture, without error
- people that don't need the best, maybe just need fast? but really are just looking for efficiency so they can do their work without running out of tokens.
Sonnet 5 doesn't meet any of those criteria. It's not the most capable, so it should either be a new standard of efficiency or speed or something. But it's not.
I guess Sonnet 5 medium maybe checks someone's box.
- mike
p.s. they did post a changelog
Changelog
Edit June 30, 2026: In the original version of this post, we included a cost-performance chart for the BrowseComp evaluation that was based on data from a simpler methodology that did not reflect the standard methodology we use for agentic search evaluations. This had the result of underestimating Sonnet 5's performance on the evaluation.
We have now updated the chart so that it matches the methodology that we used and discussed in the Sonnet 5 system card (which used a 10M token budget with compaction and programmatic tool calling). We have also updated the surrounding text.