Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I am starting to wonder how useful these benchmark still are? Aren't all of these models trained to ace these benchmarks?

    In any case they give an indication, but I am increasingly looking to real world feedback from real users. I have a solo project I (voice-to-text typing voicewink.app) and I'm using both the Claude Code and Codex coding harnesses and multiple agents doing reviews on the same code.

    This has given me real tangible results to compare on real work. Conclusion: GPT-6, Opus 5 / Fable and Grok are "top tier", with Claude good at planning and executing and Codex / Gpt-6 better at finding bugs and fixing them (but tending to overengineer), and Gemini and Kimi 3 clearly behind in capability (more so than the benchmarks suggest in my opinion).

    Any thoughts?

  • "High" to me looks like the one to use. https://artificialanalysis.ai/models/claude-opus-5-5-high

    Many benchmarks start to plateau after high, this benchmarks better than Fable, and my initial tests show it working really well.

  • Fingers are crossed on this one. I had gone back to using opus 4.8 instead of using opus 5. Simply because 4.8 is much better at remembering what it's doing and following instructions than 5. 5 often had a tendency to get halfway through solving a problem and then I would have to stop it in the middle, because it had lost its way and was going off on a tangent rather than dealing with the problem. In that respect, 4.8 was a lot more stable.
  • This continues to show that these foundational models are only slightly better than open weight models but cost around 100x as much. The history of tech is riddled with “good enough” eating “best” for lunch all day long. Unless the big labs come up with a viable business plan pronto it’s looking like AI will be no different.

    There are no prizes to be won by having the best model that’s 100x the price of something that’s good enough for 99% focuses cases.

  • Max reasoning seemed to get stuck for me too. My default is Medium which seems to work pretty well both in tight and longer running loops.
  • Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.

    Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

  • Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.
  • This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-medium

    I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

    I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.

    Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Explore Birbla archives