Discussion summary
Developers discussed the release of Grok 4.5, GPT-5.5, and Claude, noting high variability in model quality and debating whether to wait for future versions. Some emphasized immediate testing over waiting, while others expressed concerns about model restrictions and fallback behaviors.
What the discussion says
- Some users prefer to test the latest models immediately despite variability.
- Others suggest waiting for newer versions like GPT-5.6 or Sonnet.
- Concerns about model restrictions and fallback mechanisms were raised.
- One user speculated the post might be AI-generated.
“The variance in quality on these things is so high.”
“This release is a moment I will remember. The model is exactly what I want.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- The fact that the breakout previews included exactly zero gameplay is so weird to me. It shows that there was exactly zero extra effort put into anything here.by dminik
- I'd like to see the comparisons with DeepSeek, Qwen, Mimo, Kimi and GLMby Kuyawa
- I just did the tests, Mimo and GLM delivered working cubes but GLM was the only visually perfect with smooth movements and great effects.
GLM is the clear winner:
by Kuyawa - How can grok create a coding LLM at the same level as OpenAI or Anthropic when they don’t have the same amount of AI talent as the other companies by an order of magnitude? Is it really that easy to train a coding model like that?by reenorap
- Cursor
- Those two companies spend all their effort patting themselves on the back about "peer reviewed studies" and posturing immeasurables like security theatre and "trust and safety". It's really not hard to believe that Grok or Chinese models can show there is no moat.by kev009
- 1. Elon throw money on Gemini folks to get them switch ship
2. Dario is an idiot for not realising his dataset, workflow and model are going to be copied when he uses spacex datacenters
3. Grok has a special fan base that promote it everywhere they go
by throwa356262 - Love the idea, I think more complex games would show the gap in ability better.
Do it again but this time get them to make a multiplayer online Jetmen REVIVAL game. Online play is key, because it's very complex. Jetmen is a good game for this since it has physics and customization that's complex enough but still simple.
by singingtoday - Why not wait one more day for GPT-5.6?by paxys
- And why not Sonnet?by fluidcruft
- Well, I'm probably not on the list of special people who will get to see GPT-5.6 Terra.by trollbridge
- I worry that GPT 5.6 will be heavily restricted and have the same feature to fallback to another model like Claude fable 5 does all too often. That fallback shenanigans mess up actual benchmarks and I don't like it.by acters
- That will be in the Part 2 article.by m4rkuskk
- Also throw in GLM 5.2 for good measureby faitswulff
- If we wait for the next models, we will never test anything because there will always be another model. Like the Ai Scotsman:
> "Nay, laddie, that’s no’ the real AI Scotsman! He’s grander still! More powerful! Just wait for the next model!"
by foxfired - > The receipts: speed and cost
I don't get why cost per reply is at all relevant here?
Why do so few who attempt comparisons actually compare dollars per task.
by dirteater_ - Tokens per task might be the better choiceby baxtr
- I think these were all one shots, so it was 1 reply per task?by NewJazz
- If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/
I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).
by sgk284 - Impressive, specially the amount of models used for the comparison.by coopykins
- This is impressive and, I think, complements well whatever benchmark is the hottest right now.by Otterly99
- Really nice site! From your experience, what’s your go-to model for nice storefronts?
- I want this with smaller models as well like Gemma 4 or Qwen 3.6by kristopolous
- I have not used grok 4.5 yet, but the other pictures match my experience doing anything graphical with the other models that it cracks me up. gpt 5.5 has no design sense whatsoever. It cannot even make terminal output not look terrible. I've asked it to use colors and formatting in various ways and got goofy randomly colored output. opus 4.7 and later seemed to have an inuitive design sense by comparison - 2d or 3d. Fabel 5 is just rock solid.
Yes, subjective. But it matches my repeated experiences with these models for what it is worth.
- I get better results with Opus than Fable 5 on various oneshots (including our old friend of "generate an SVG of a pelican riding a bicycle"). (Opus is also far and away better than pretty much any other SOTA or near-SOTA model.)by trollbridge
- So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?by jeffgreco
- GPT was the worst on the Rubik's cubeby singingtoday
- I am 99% sure the post was written by AIby mlmonkey
- > Role reversal, two figures, one file.
> “snappy stylist”
Funny you can tell its slop just by this
by HeartStrings - I'll actually put in spelling errors and grammar mistakes intentionally these days to show humanby kristopolous