

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I've had much better experiences with the "budget" tier models than this test suggests I should. I rarely even consider the higher models due to price/use. Perhaps I'm just more vigilant about spec'ing my prompts out before submitting them? Am I just doing too pleb of work? I'm not trying to write custom cuda kernels or hardware integration.by citizenpaul
- What happened in July 2025 and January 2026? If that can't be explained in a concise way this entire dataset is vibes at best.by tcdent
- Why are models better than agents, isn't it supposed to be the opposite? I don't understand the difference and what you are measuring.by goldenarm
- Consistently impressed with the performance of Grok, given they started so much later than everyone else and, in some ways, are less well funded compared to OpenAI and Anthropic. I wonder how much of this is just luck, name recognition, or management style. Marc Andreesen likes to talk about how Elon companies have a unique engineering-heavy management structure, which contrasts the research heavy cultures of OpenAI and other labs, and it makes me wonder if that sort of thing could be behind their relative success.by guywithahat
- Call me biased, but if Grok 4.5 is above GPT-5.6 Sol, I don’t trust this benchmark.by stared
- Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels.by dia80
- They are all different problems for the different languages. I was hoping this was a benchmark that attempted to see which languages were more efficient to use with which models.by spullara
- What does it mean when Fable 5 is 1st place and Opus 5 is 3rd place, while Claude code is 7th place? Which model and effort is used for Claude in 7th place, compared to 1st and 3rd?by sathish316