Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I've had much better experiences with the "budget" tier models than this test suggests I should. I rarely even consider the higher models due to price/use. Perhaps I'm just more vigilant about spec'ing my prompts out before submitting them? Am I just doing too pleb of work? I'm not trying to write custom cuda kernels or hardware integration.
  • What happened in July 2025 and January 2026? If that can't be explained in a concise way this entire dataset is vibes at best.
  • Why are models better than agents, isn't it supposed to be the opposite? I don't understand the difference and what you are measuring.
  • Consistently impressed with the performance of Grok, given they started so much later than everyone else and, in some ways, are less well funded compared to OpenAI and Anthropic. I wonder how much of this is just luck, name recognition, or management style. Marc Andreesen likes to talk about how Elon companies have a unique engineering-heavy management structure, which contrasts the research heavy cultures of OpenAI and other labs, and it makes me wonder if that sort of thing could be behind their relative success.
  • Call me biased, but if Grok 4.5 is above GPT-5.6 Sol, I don’t trust this benchmark.
  • Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels.
  • They are all different problems for the different languages. I was hoping this was a benchmark that attempted to see which languages were more efficient to use with which models.
  • What does it mean when Fable 5 is 1st place and Opus 5 is 3rd place, while Claude code is 7th place? Which model and effort is used for Claude in 7th place, compared to 1st and 3rd?

Explore Birbla archives