Discussion summary

Developers discussed the release of Grok 4.5, GPT-5.5, and Claude, noting high variability in model quality and debating whether to wait for future versions. Some emphasized immediate testing over waiting, while others expressed concerns about model restrictions and fallback behaviors.

What the discussion says

  • Some users prefer to test the latest models immediately despite variability.
  • Others suggest waiting for newer versions like GPT-5.6 or Sonnet.
  • Concerns about model restrictions and fallback mechanisms were raised.
  • One user speculated the post might be AI-generated.
“The variance in quality on these things is so high.”
— RickS
“This release is a moment I will remember. The model is exactly what I want.”
— maxdo

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Why not wait one more day for GPT-5.6?
  • > The receipts: speed and cost

    I don't get why cost per reply is at all relevant here?

    Why do so few who attempt comparisons actually compare dollars per task.

  • If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/

    I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).

  • I have not used grok 4.5 yet, but the other pictures match my experience doing anything graphical with the other models that it cracks me up. gpt 5.5 has no design sense whatsoever. It cannot even make terminal output not look terrible. I've asked it to use colors and formatting in various ways and got goofy randomly colored output. opus 4.7 and later seemed to have an inuitive design sense by comparison - 2d or 3d. Fabel 5 is just rock solid.

    Yes, subjective. But it matches my repeated experiences with these models for what it is worth.

  • So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?
  • I am 99% sure the post was written by AI
  • Half year ago I tried to use Codex, Claude and Gemini build the same scripts to automate various things on my machine. Claude was the clear winner back then, making the most reasonable assumptions, presenting results in the easiest-to-read format, writing runnable script with minimum dependency. Half year later I think Codex and Claude models have both advanced a lot, but Gemini is still lackluster. Gemini could catch problems when reviewing Claude/Codex's design plans and code, but it's hard to make Gemini make complex plans or implement complex code by itself.
  • I tried to one-shot the first test (the Rubik's Cube test) with LucidQuery's Swift model, to test it, as there are not much benchmarks about it and that they brag a lot about it, and I was pleasantly surprised to see it achieving a result similar to Grok 4.5 but in one shot (there is the same issue that if you scramble twice the solve button does not work anymore, but it got it in one shot).

    Though it crunched most of the free quota, 47111 tokens, so I couldn't make multiple attempts.

Explore Birbla archives