

Discussion summary
Discussions about AI benchmarks and model performance, with concerns about approaching 100% accuracy and the implications. Updates on Opus 5 and comparisons to other models like Fable are also mentioned.
What the discussion says
- Some see increasing scores as a sign of progress, questioning what near-perfect accuracy means.
- Concerns about the transparency of costs and the distortion in cost graphs.
- Debate on the significance of benchmarks and their real-world applicability.
“Do the ever increasing scores mean models will soon approach 100%? And what would that even mean?”
“It's not Fable, but I'll take it.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Seems to be another great incremental update to the workhorse, nice!
I've been using Sonnet instead of Opus for almost all coding tasks for a while now. A little elbow grease to break down tasks and you can spend a lot less money for just about the same output quality.
- Important to note: "Sonnet 5 is an upgrade to Sonnet 4.6, but it uses an updated tokenizer that changes how the model processes text to improve performance (this is similar to the tokenizer change we introduced with Claude Opus 4.7). The tradeoff is that the same input can map to more tokens: roughly 1.0–1.35× depending on the content type. The introductory pricing is set so that the transition to Sonnet 5 is roughly cost-neutral."by m3h
- Wonder if the whole cyber paranoia leads to their models ultimately generating less secure code. After all, if it has the ability to generate safe code, it would imply that it knows something about cybersecurity, which could surely be used to hack all the banks in the world.by Sol-
- Claude Sonnet 5 itself described its pelican as looking like a goose:
> Illustration of a white goose riding a bicycle, with one wing extended forward to grip the handlebar, set against a plain white background with a brown ground line.
by simonw - I just tested it on my benchmarks[0], it's GLM-5.2 level, at 2x cost, but also 2x faster.
Weak spots (categories it fails):
[0]: https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...- Trivia — 0/3 - basically not much built-in knowledge - Combined tool-calling tasks — score 45/100, sometimes makes invalid tool calls - Puzzle Solving — score 77, flubs carwash-like testsby XCSme - Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.
I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have found that the more models are optimized for fully agentic development, the worse they get at assisted development and often start doing too much despite very strict/specific instructions.
I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap.
by microtonal - Wow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters.
From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5
As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberGym"
by conradkay - I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort.
Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.
by Jcampuzano2